The solution uses Ollama as the local LLM runtime, TranslateGemma 4B as the primary translation model, and Python as the orchestration layer.
The approach also evaluates two different models: Qwen3 4B and TranslateGemma 4B.

1. Problem Statement
The application already contains approximately 12,000 multiple-choice questions (MCQs) stored in PostgreSQL. Each MCQ consists of a question, four options, the correct answer, and an explanation.A typical record looks like this:
{
"question": "Which of the following statements are true about India?\n
\n1. India is the fifth largest country in the world.
\n2. It occupies about 2.4 percent of the total area of the lithosphere.
\n3. Whole of India lies in the torrid zone.
\n4. 82┬░30' East Meridian and hemisphere is used to determine Indian Standard Time",
"optionA": "1 and 2",
"optionB": "2 and 3",
"optionC": "1 and 3",
"optionD": "2 and 4",
"explanation": "India's total geographical area is approximately 3.28 million square kilometers, which accounts for about 2.4% of the world's total geographical area.\n
\nThe 82┬░30' E longitude is the Standard Meridian of India, which passes through the center of the country and is used as the reference meridian for Indian Standard Time (IST).\n
\nIndia is the seventh-largest country in the world by area. The countries larger than India are Russia, Canada, China, the United States, Brazil, and Australia.\n
\nThe Tropic of Cancer (23┬░30' N latitude) passes through the middle of the country, dividing it into almost two equal parts. The area south of the Tropic of Cancer lies in the torrid (tropical) zone, while the area north of it lies in the temperate (subtropical) zone.",
"answer": "D",
"quiz": {
"name": "India: Area, Latitude, Tropic of Cancer & Borders",
"subCategory": {
"name": "Indian Geography",
"category": {
"name": "General Studies (GS) PYQs"
}
}
},
"tags": [
"Indian Geography"
],
"uniqueCode": "X1y0ljrNx",
"slug": "which-of-the-following-statements-are-true-about-india-br"
}
The requirement is to generate a Hindi version while preserving the database identity and MCQ structure.
Why local LLM?
When an application contains thousands of educational questions, translating them manually is expensive and difficult to maintain.
Sending every question to a commercial LLM API such as ChatGPT or Gemini can also become expensive when the dataset contains thousands of questions, options, and detailed explanations.
A better approach for a large batch-translation workload is to run an open translation model locally.
The model runs on the developer's own machine, the English content remains local, and the translation cost is effectively zero after the initial setup.
2. Model Candidates
I tried two models during the evaluation.The first was Qwen3 4B. Qwen3 is a general-purpose multilingual LLM rather than a model dedicated exclusively to translation.
The Q4_K_M version used in the experiment is approximately 3.5 GB and can run locally through Ollama. Ollama exposes Qwen3's thinking capability, which can be disabled through the API using "think": false.
The second model was TranslateGemma 4B. It is specifically designed for translation and is based on Gemma 3.
The Ollama distribution provides 4B, 12B, and 27B variants, with the 4B version approximately 3.3 GB. TranslateGemma supports translation across 55 languages and has a specific prompt format intended for translation tasks.
2.1 Qwen3 4B Experiment
The first experiment used Qwen3 4B through Ollama. The model was installed using:ollama pull qwen3:4b-q4_K_M
The model could then be executed locally:
ollama run qwen3:4b-q4_K_M
% ollama ps
NAME ID SIZE PROCESSOR CONTEXT UNTIL
qwen3:4b-q4_K_M 2bfd38a7daaf 3.3 GB 100% GPU 4096 4 minutes from now
The model showed approximately 4B parameters with Q4_K_M quantization.
For the API-based translation test, the following request was used:
request-qwen.json
{
"model": "qwen3:4b-q4_K_M",
"messages": [
{
"role": "user",
"content": "You are a professional English (en) to Hindi (hi) translator. Translate the following MCQ into standard, natural Hindi suitable for Indian competitive-exam questions.\n\nRules:\n1. Return ONLY the Hindi translation.\n2. Do not add explanations, commentary, markdown, or code fences.\n3. Preserve all HTML tags exactly, including , , and
.\n4. Preserve numbers, dates, percentages, degrees, coordinates, symbols, and units.\n5. Use commonly accepted Hindi names for Indian places, geographical terms, historical terms, institutions, and technical terminology.\n6. Do not mechanically transliterate English words when an established Hindi term exists.\n7. Do not change the factual meaning.\n8. Translate the question, optionA, optionB, optionC, optionD, and explanation.\n\nQuestion:\nWhich of the following statements are true about India?\n
\n1. India is the fifth largest country in the world.\n2. It occupies about 2.4 percent of the total area of the lithosphere.\n3. Whole of India lies in the torrid zone.\n4. 82┬░30' East Meridian and hemisphere is used to determine Indian Standard Time\n\nOption A:\n1 and 2\n\nOption B:\n2 and 3\n\nOption C:\n1 and 3\n\nOption D:\n2 and 4\n\nExplanation:\nIndia's total geographical area is approximately 3.28 million square kilometers, which accounts for about 2.4% of the world's total geographical area.\n
\nThe 82┬░30' E longitude is the Standard Meridian of India, which passes through the center of the country and is used as the reference meridian for Indian Standard Time (IST).\n
\nIndia is the seventh-largest country in the world by area. The countries larger than India are Russia, Canada, China, the United States, Brazil, and Australia.\n
\nThe Tropic of Cancer (23┬░30' N latitude) passes through the middle of the country, dividing it into almost two equal parts. The area south of the Tropic of Cancer lies in the torrid (tropical) zone, while the area north of it lies in the temperate (subtropical) zone."
}
],
"think": false,
"stream": false
}
curl http://localhost:11434/api/chat \
-H "Content-Type: application/json" \
-d @request-qwen.json \
-o response-qwen.json
% Total % Received % Xferd Average Speed Time Time Time Current
Dload Upload Total Spent Left Speed
100 5263 0 3063 100 2200 89 64 0:00:34 0:00:34 --:--:-- 786
response-qwen.json
{
"model": "qwen3:4b-q4_K_M",
"created_at": "2026-08-11T16:58:21.816271Z",
"message": {
"role": "assistant",
"content": "рдЕрдВрддрд░реНрд░рд╛рд╖реНрдЯреНрд░реАрдп рдШрдЯрдирд╛рдУрдВ рдХреЗ рдмрд╛рд░реЗ рдореЗрдВ рдирд┐рдореНрдирд▓рд┐рдЦрд┐рдд рдХрдердиреЛрдВ рдореЗрдВ рд╕реЗ рдХреМрди-рд╕рд╛ рд╕рддреНрдп рд╣реИ?\n
\n1. рднрд╛рд░рдд рд╡рд┐рд╢реНрд╡ рдХреЗ рдкрд╛рдВрдЪрд╡реЗ рд╕рдмрд╕реЗ рдмрдбрд╝рд╛ рджреЗрд╢ рд╣реИред\n2. рдпрд╣ рдкреГрдереНрд╡реА рдХреЗ рдХреБрд▓ рдХреНрд╖реЗрддреНрд░рдлрд▓ рдХреЗ рд▓рд┐рдереЛрд╕реНрдлреАрд░реЗ рдХреЗ рд▓рдЧрднрдЧ 2.4 рдкреНрд░рддрд┐рд╢рдд рдХрд╛ рдЖрдВрдХрдбрд╝рд╛ рджреЗрддрд╛ рд╣реИред\n3. рднрд╛рд░рдд рдХреЗ рдкреВрд░реЗ рдХреНрд╖реЗрддреНрд░ рд╢рд╛рд░реАрд░рд┐рдХ рдЬреЛрди рдореЗрдВ рд░рд╣рддрд╛ рд╣реИред\n4. 82┬░30' рдкреВрд░реНрд╡ рдЕрдХреНрд╖рд╛рдВрд╢ рдФрд░ рдЖрдзрд┐рдХрд╛рд░рд┐рдХ рдЕрд░реНрдзрдЧреЛрд▓рд╛рд░реНрджреНрдз рднрд╛рд░рддреАрдп рд╕рдордп рдХреЗ рдирд┐рд░реНрдзрд╛рд░рдг рдХреЗ рд▓рд┐рдП рдЙрдкрдпреЛрдЧ рдХрд┐рдпрд╛ рдЬрд╛рддрд╛ рд╣реИ\n\nрд╡рд┐рдХрд▓реНрдк A:\n1 рдФрд░ 2\n\nрд╡рд┐рдХрд▓реНрдк B:\n2 рдФрд░ 3\n\nрд╡рд┐рдХрд▓реНрдк C:\n1 рдФрд░ 3\n\nрд╡рд┐рдХрд▓реНрдк D:\n2 рдФрд░ 4\n\nрд╕реНрдкрд╖реНрдЯреАрдХрд░рдг:\nрднрд╛рд░рдд рдХреЗ рдХреБрд▓ рднреВрдЧреЛрд▓реАрдп рдХреНрд╖реЗрддреНрд░ рд▓рдЧрднрдЧ 3.28 рдорд┐рд▓рд┐рдпрди рд╡рд░реНрдЧ рдХрд┐рд▓реЛрдореАрдЯрд░ рд╣реИ, рдЬреЛ рд╡рд┐рд╢реНрд╡ рдХреЗ рдХреБрд▓ рднреВрдЧреЛрд▓реАрдп рдХреНрд╖реЗрддреНрд░ рдХреЗ рд▓рдЧрднрдЧ 2.4% рдХрд╛ рдЖрдВрдХрдбрд╝рд╛ рджреЗрддрд╛ рд╣реИред\n
\n82┬░30' E рдЕрдХреНрд╖рд╛рдВрд╢ рднрд╛рд░рдд рдХрд╛ рд╕рд╛рдорд╛рдиреНрдп рдЕрдХреНрд╖рд╛рдВрд╢ рд╣реИ, рдЬреЛ рджреЗрд╢ рдХреЗ рдХреЗрдВрджреНрд░ рдореЗрдВ рдЧреБрдЬрд░рддрд╛ рд╣реИ рдФрд░ рднрд╛рд░рддреАрдп рд╕рдордп (IST) рдХреЗ рд▓рд┐рдП рд╕рдВрджрд░реНрдн рдЕрдХреНрд╖рд╛рдВрд╢ рдХреЗ рд░реВрдк рдореЗрдВ рдЙрдкрдпреЛрдЧ рдХрд┐рдпрд╛ рдЬрд╛рддрд╛ рд╣реИред\n
\nрднрд╛рд░рдд рд╡рд┐рд╢реНрд╡ рдореЗрдВ рднреВрдЧреЛрд▓реАрдп рдХреНрд╖реЗрддреНрд░ рдХреЗ рдЖрдзрд╛рд░ рдкрд░ рд╕рд╛рддрд╡реЗрдВ рд╕рдмрд╕реЗ рдмрдбрд╝рд╛ рджреЗрд╢ рд╣реИред рднрд╛рд░рдд рдХреЗ рдмрдбрд╝реЗ рджреЗрд╢ рдЗрд╕рдХреЗ рдЕрддрд┐рд░рд┐рдХреНрдд рд░реВрд╕, рдХреИрдирд╛рдбрд╛, рдЪреАрди, рдЕрдореЗрд░рд┐рдХрд╛, рдмреНрд░рд╛рдЬреАрд▓ рдФрд░ рдСрд╕реНрдЯреНрд░реЗрд▓рд┐рдпрд╛ рд╣реИрдВред\n
\nрдХреИрдВрд╕рд░ рд╡рд┐рдкрддреНрддрд┐ (23┬░30' рдЙрддреНрддрд░ рдЕрдХреНрд╖рд╛рдВрд╢) рджреЗрд╢ рдХреЗ рдХреЗрдВрджреНрд░ рдореЗрдВ рдЧреБрдЬрд░рддрд╛ рд╣реИ, рдЬреЛ рджреЗрд╢ рдХреЛ рд▓рдЧрднрдЧ рджреЛ рдмрд░рд╛рдмрд░ рднрд╛рдЧреЛрдВ рдореЗрдВ рд╡рд┐рднрд╛рдЬрд┐рдд рдХрд░рддрд╛ рд╣реИред рдХреИрдВрд╕рд░ рд╡рд┐рдкрддреНрддрд┐ рдХреЗ рджрдХреНрд╖рд┐рдг рдореЗрдВ рд╕реНрдерд┐рдд рдХреНрд╖реЗрддреНрд░ рд╢рд╛рд░реАрд░рд┐рдХ (рддрд╛рдкрдорд╛рди) рдЬреЛрди рдореЗрдВ рд░рд╣рддрд╛ рд╣реИ, рдЬрдмрдХрд┐ рдЗрд╕рдХреЗ рдЙрддреНрддрд░ рдореЗрдВ рд╕реНрдерд┐рдд рдХреНрд╖реЗрддреНрд░ рд╢реАрдд рдЬреЛрди (рд╕рдмрдЯреНрд░реЛрдкрд┐рдХ) рдореЗрдВ рд░рд╣рддрд╛ рд╣реИред"
},
"done": true,
"done_reason": "stop",
"total_duration": 36274259750,
"load_duration": 186300625,
"prompt_eval_count": 515,
"prompt_eval_duration": 54102000,
"eval_count": 1071,
"eval_duration": 35877791000
}
There are several serious translation errors, such as: India тЖТ рдЕрдВрддрд░реНрд░рд╛рд╖реНрдЯреНрд░реАрдп рдШрдЯрдирд╛рдУрдВ. The original question was about India, but Qwen changed the subject completely.
Qwen3 therefore worked technically, but the experiment showed that a general-purpose LLM was not necessarily the best choice for a large-scale English-to-Hindi translation workload.
2.2 TranslateGemma 4B
The second model was TranslateGemma 4B.This model was immediately more interesting because it is specifically designed for translation while still being small enough to run locally through Ollama.
The official Ollama distribution lists the 4B version at approximately 3.0 GB and supports 55 languages.
Installation is extremely simple:
ollama pull translategemma:4b
The model can then be run locally:
ollama run translategemma:4b
% ollama ps
NAME ID SIZE PROCESSOR CONTEXT UNTIL
translategemma:4b c49d986b0764 2.9 GB 100% GPU 4096 3 minutes from now
The following test was performed using exactly the same MCQ used for Qwen3:
request-translategemma.json
{
"model": "translategemma:4b",
"messages": [
{
"role": "user",
"content": "You are a professional English (en) to Hindi (hi) translator. Translate the following MCQ into standard, natural Hindi suitable for Indian competitive-exam questions.\n\nRules:\n1. Return ONLY the Hindi translation.\n2. Do not add explanations, commentary, markdown, or code fences.\n3. Preserve all HTML tags exactly, including , , and
.\n4. Preserve numbers, dates, percentages, degrees, coordinates, symbols, and units.\n5. Use commonly accepted Hindi names for Indian places, geographical terms, historical terms, institutions, and technical terminology.\n6. Do not mechanically transliterate English words when an established Hindi term exists.\n7. Do not change the factual meaning.\n8. Translate the question, optionA, optionB, optionC, optionD, and explanation.\n\nQuestion:\nWhich of the following statements are true about India?\n
\n1. India is the fifth largest country in the world.\n2. It occupies about 2.4 percent of the total area of the lithosphere.\n3. Whole of India lies in the torrid zone.\n4. 82┬░30' East Meridian and hemisphere is used to determine Indian Standard Time\n\nOption A:\n1 and 2\n\nOption B:\n2 and 3\n\nOption C:\n1 and 3\n\nOption D:\n2 and 4\n\nExplanation:\nIndia's total geographical area is approximately 3.28 million square kilometers, which accounts for about 2.4% of the world's total geographical area.\n
\nThe 82┬░30' E longitude is the Standard Meridian of India, which passes through the center of the country and is used as the reference meridian for Indian Standard Time (IST).\n
\nIndia is the seventh-largest country in the world by area. The countries larger than India are Russia, Canada, China, the United States, Brazil, and Australia.\n
\nThe Tropic of Cancer (23┬░30' N latitude) passes through the middle of the country, dividing it into almost two equal parts. The area south of the Tropic of Cancer lies in the torrid (tropical) zone, while the area north of it lies in the temperate (subtropical) zone."
}
],
"stream": false
}
curl http://localhost:11434/api/chat \
-H "Content-Type: application/json" \
-d @request-translategemma.json \
-o response-translategemma.json
% Total % Received % Xferd Average Speed Time Time Time Current
Dload Upload Total Spent Left Speed
100 5493 0 3308 100 2185 247 163 0:00:13 0:00:13 --:--:-- 805
response-translategemma.json
{
"model": "translategemma:4b",
"created_at": "2026-08-11T17:09:24.355449Z",
"message": {
"role": "assistant",
"content": "рдкреНрд░рд╢реНрди:\nрдирд┐рдореНрдирд▓рд┐рдЦрд┐рдд рдореЗрдВ рд╕реЗ рднрд╛рд░рдд рдХреЗ рдмрд╛рд░реЗ рдореЗрдВ рдХреМрди рд╕реА рдХрдерди рд╕рддреНрдп рд╣реИрдВ?\n
\n1. рднрд╛рд░рдд рджреБрдирд┐рдпрд╛ рдХрд╛ рдкрд╛рдВрдЪрд╡рд╛рдВ рд╕рдмрд╕реЗ рдмрдбрд╝рд╛ рджреЗрд╢ рд╣реИред\n2. рдпрд╣ рдХреБрд▓ рднреВ-рднрд╛рдЧ рдХреНрд╖реЗрддреНрд░ рдХрд╛ рд▓рдЧрднрдЧ 2.4 рдкреНрд░рддрд┐рд╢рдд рднрд╛рдЧ рд╣реИред\n3. рдкреВрд░реЗ рднрд╛рд░рдд рдХрд╛ рдХреНрд╖реЗрддреНрд░ рдЙрд╖реНрдгрдХрдЯрд┐рдмрдВрдзреАрдп рдХреНрд╖реЗрддреНрд░ рдореЗрдВ рд╕реНрдерд┐рдд рд╣реИред\n4. 82┬░30' рдкреВрд░реНрд╡ рджреЗрд╢рд╛рдВрддрд░ рдФрд░ рдЕрд░реНрдз-рдЧреЛрд▓рд╛ рдХрд╛ рдЙрдкрдпреЛрдЧ рднрд╛рд░рддреАрдп рдорд╛рдирдХ рд╕рдордп рдирд┐рд░реНрдзрд╛рд░рд┐рдд рдХрд░рдиреЗ рдХреЗ рд▓рд┐рдП рдХрд┐рдпрд╛ рдЬрд╛рддрд╛ рд╣реИред\n\nрд╡рд┐рдХрд▓реНрдк A:\n1 рдФрд░ 2\n\nрд╡рд┐рдХрд▓реНрдк B:\n2 рдФрд░ 3\n\nрд╡рд┐рдХрд▓реНрдк C:\n1 рдФрд░ 3\n\nрд╡рд┐рдХрд▓реНрдк D:\n2 рдФрд░ 4\n\nрд╡реНрдпрд╛рдЦреНрдпрд╛:\nрднрд╛рд░рдд рдХрд╛ рдХреБрд▓ рднреМрдЧреЛрд▓рд┐рдХ рдХреНрд╖реЗрддреНрд░рдлрд▓ рд▓рдЧрднрдЧ 3.28 рдорд┐рд▓рд┐рдпрди рд╡рд░реНрдЧ рдХрд┐рд▓реЛрдореАрдЯрд░ рд╣реИ, рдЬреЛ рд╡рд┐рд╢реНрд╡ рдХреЗ рдХреБрд▓ рднреМрдЧреЛрд▓рд┐рдХ рдХреНрд╖реЗрддреНрд░рдлрд▓ рдХрд╛ рд▓рдЧрднрдЧ 2.4% рд╣реИред\n
\n82┬░30' рдкреВрд░реНрд╡ рдЕрдХреНрд╖рд╛рдВрд╢ рднрд╛рд░рдд рдХрд╛ рдорд╛рдирдХ рджреЗрд╢рд╛рдВрддрд░ рд╣реИ, рдЬреЛ рджреЗрд╢ рдХреЗ рдХреЗрдВрджреНрд░ рд╕реЗ рд╣реЛрдХрд░ рдЧреБрдЬрд░рддрд╛ рд╣реИ рдФрд░ рдЗрд╕рдХрд╛ рдЙрдкрдпреЛрдЧ рднрд╛рд░рддреАрдп рдорд╛рдирдХ рд╕рдордп (IST) рдХреЗ рд▓рд┐рдП рд╕рдВрджрд░реНрдн рджреЗрд╢рд╛рдВрддрд░ рдХреЗ рд░реВрдк рдореЗрдВ рдХрд┐рдпрд╛ рдЬрд╛рддрд╛ рд╣реИред\n
\nрднрд╛рд░рдд рдХреНрд╖реЗрддреНрд░рдлрд▓ рдХреЗ рд╣рд┐рд╕рд╛рдм рд╕реЗ рджреБрдирд┐рдпрд╛ рдХрд╛ рд╕рд╛рддрд╡рд╛рдВ рд╕рдмрд╕реЗ рдмрдбрд╝рд╛ рджреЗрд╢ рд╣реИред рд░реВрд╕, рдХрдирд╛рдбрд╛, рдЪреАрди, рд╕рдВрдпреБрдХреНрдд рд░рд╛рдЬреНрдп рдЕрдореЗрд░рд┐рдХрд╛, рдмреНрд░рд╛рдЬреАрд▓ рдФрд░ рдСрд╕реНрдЯреНрд░реЗрд▓рд┐рдпрд╛ рдЬреИрд╕реЗ рджреЗрд╢реЛрдВ рдХрд╛ рдХреНрд╖реЗрддреНрд░рдлрд▓ рднрд╛рд░рдд рд╕реЗ рдЕрдзрд┐рдХ рд╣реИред\n
\nрд╡реГрддреНрдд рдордзреНрдпрд╛рдВрд╢ (23┬░30' рдЙрддреНрддрд░ рдЕрдХреНрд╖рд╛рдВрд╢) рдкреВрд░реЗ рджреЗрд╢ рд╕реЗ рд╣реЛрдХрд░ рдЧреБрдЬрд░рддрд╛ рд╣реИ рдФрд░ рдЗрд╕реЗ рджреЛ рд▓рдЧрднрдЧ рд╕рдорд╛рди рднрд╛рдЧреЛрдВ рдореЗрдВ рд╡рд┐рднрд╛рдЬрд┐рдд рдХрд░рддрд╛ рд╣реИред рд╡реГрддреНрдд рдордзреНрдпрд╛рдВрд╢ рдХреЗ рджрдХреНрд╖рд┐рдг рдореЗрдВ рд╕реНрдерд┐рдд рдХреНрд╖реЗрддреНрд░ рдЙрд╖реНрдгрдХрдЯрд┐рдмрдВрдзреАрдп (рдЯреЛрд░рд┐рдб) рдХреНрд╖реЗрддреНрд░ рдореЗрдВ рд╣реИ, рдЬрдмрдХрд┐ рдЗрд╕рдХреЗ рдЙрддреНрддрд░реА рднрд╛рдЧ рдореЗрдВ рдЙрдк-рдЙрд╖реНрдгрдХрдЯрд┐рдмрдВрдзреАрдп (рдЯреЗрдореНрдкрд░реЗрдЯ) рдХреНрд╖реЗрддреНрд░ рд╣реИред"
},
"done": true,
"done_reason": "stop",
"total_duration": 13274667417,
"load_duration": 346484209,
"prompt_eval_count": 534,
"prompt_eval_duration": 521600000,
"eval_count": 386,
"eval_duration": 12314585000
}
TranslateGemma's output is substantially better than the Qwen3 output.
TranslateGemma is not perfect, but its behavior was considerably more appropriate for this particular translation workload.
3. Translation
After evaluating the two models, I used TranslateGemma 4B for the actual translation process.Since the application already stores all MCQs in PostgreSQL, there was no need to create a separate translation database or duplicate the existing records.
The approach was to extend the existing MCQ table with Hindi columns and a translated flag.
The original English content remains unchanged, while the Hindi translation is stored alongside it in the same database record.
The database structure was extended using:
ALTER TABLE public.questions
ADD COLUMN question_hindi TEXT,
ADD COLUMN option_a_hindi TEXT,
ADD COLUMN option_b_hindi TEXT,
ADD COLUMN option_c_hindi TEXT,
ADD COLUMN option_d_hindi TEXT,
ADD COLUMN explanation_hindi TEXT,
ADD COLUMN translated BOOLEAN DEFAULT FALSE;
The translated column is important because the translation process is not necessarily completed in a single run.
With approximately 12,000 MCQs, the process can take several hours depending on the model, hardware, and length of each explanation.
Instead of processing all records every time the Python script starts, the script reads only records where translated = false.
After a successful translation, it writes the Hindi values back to the same database record and changes the flag to true.
The Python application communicates with Ollama through its local HTTP API.
Because Ollama is running on the same machine, the request is sent to http://localhost:11434 and does not require an external LLM API.
The Python dependencies required for this implementation are:
pip install psycopg2-binary requests
The database connection can be configured using environment variables rather than hard-coding database credentials in the application.
import os
import psycopg2
import requests
DB_HOST = os.getenv("DB_HOST", "localhost")
DB_PORT = os.getenv("DB_PORT", "5432")
DB_NAME = os.getenv("DB_NAME", "mcq_db")
DB_USER = os.getenv("DB_USER", "postgres")
DB_PASSWORD = os.getenv("DB_PASSWORD", "")
OLLAMA_URL = "http://localhost:11434/api/chat"
MODEL = "translategemma:4b"
def get_connection():
return psycopg2.connect(
host=DB_HOST,
port=DB_PORT,
database=DB_NAME,
user=DB_USER,
password=DB_PASSWORD
)
The translation request contains the MCQ fields that need to be translated.
The database identity fields such as the record ID, answer, quiz information, tags, and slug are not sent as translatable content.
The prompt explicitly tells TranslateGemma to return only the translated content and preserve HTML tags, numbers, dates, percentages, symbols, and other important formatting.
def translate_mcq(mcq):
prompt = f"""
You are a professional English (en) to Hindi (hi) translator.
Translate the following MCQ into standard, natural Hindi
suitable for Indian competitive-exam questions.
Rules:
1. Return ONLY JSON matching the required schema.
2. Do not add markdown or explanations.
3. Preserve all HTML tags exactly, including , , and
.
4. Preserve numbers, dates, percentages, degrees, coordinates,
symbols, and units.
5. Use commonly accepted Hindi terminology.
6. Do not mechanically transliterate English words when an
established Hindi term exists.
7. Do not change the factual meaning.
8. Translate ONLY the values of question, optionA, optionB,
optionC, optionD, and explanation.
9. The JSON property names MUST remain exactly:
question, optionA, optionB, optionC, optionD, explanation.
Input:
{{
"question": {mcq["question"]!r},
"optionA": {mcq["option_a"]!r},
"optionB": {mcq["option_b"]!r},
"optionC": {mcq["option_c"]!r},
"optionD": {mcq["option_d"]!r},
"explanation": {mcq["explanation"]!r}
}}
"""
schema = {
"type": "object",
"properties": {
"question": {"type": "string"},
"optionA": {"type": "string"},
"optionB": {"type": "string"},
"optionC": {"type": "string"},
"optionD": {"type": "string"},
"explanation": {"type": "string"}
},
"required": [
"question",
"optionA",
"optionB",
"optionC",
"optionD",
"explanation"
]
}
response = requests.post(
OLLAMA_URL,
json={
"model": MODEL,
"messages": [
{
"role": "user",
"content": prompt
}
],
"stream": False,
"format": schema
},
timeout=REQUEST_TIMEOUT
)
response.raise_for_status()
content = response.json()["message"]["content"]
print("Raw model response:")
print(content)
return json.loads(content)
The translation script then reads a batch of untranslated MCQs from PostgreSQL. For each record, it sends the content to TranslateGemma through Ollama and receives the Hindi translation.
After receiving the response, the script updates the Hindi columns and marks the record as translated.
import json
def get_untranslated_mcqs(connection, batch_size=10):
with connection.cursor() as cursor:
cursor.execute(
"""
SELECT
id,
question,
option_a,
option_b,
option_c,
option_d,
explanation
FROM public.questions
WHERE translated = FALSE
ORDER BY id
LIMIT %s
""",
(batch_size,)
)
return cursor.fetchall()
def update_translation(connection, mcq_id, translation):
with connection.cursor() as cursor:
cursor.execute(
"""
UPDATE public.questions
SET
question_hindi = %s,
option_a_hindi = %s,
option_b_hindi = %s,
option_c_hindi = %s,
option_d_hindi = %s,
explanation_hindi = %s,
translated = TRUE
WHERE id = %s
""",
(
translation["question"],
translation["optionA"],
translation["optionB"],
translation["optionC"],
translation["optionD"],
translation["explanation"],
mcq_id
)
)
connection.commit()
The complete processing loop can then continuously fetch untranslated records until there are no records left.
def translate_all():
connection = get_connection()
try:
while True:
rows = get_untranslated_mcqs(
connection,
batch_size=10
)
if not rows:
print("Translation completed.")
break
for row in rows:
mcq = {
"id": row[0],
"question": row[1],
"option_a": row[2],
"option_b": row[3],
"option_c": row[4],
"option_d": row[5],
"explanation": row[6]
}
print(f"Translating MCQ: {mcq['id']}")
try:
result = translate_mcq(mcq)
translation = json.loads(result)
update_translation(
connection,
mcq["id"],
translation
)
print(
f"MCQ {mcq['id']} translated successfully."
)
except Exception as e:
connection.rollback()
print(
f"Failed to translate MCQ "
f"{mcq['id']}: {e}"
)
finally:
connection.close()
if __name__ == "__main__":
translate_all()
The important part of this design is that the database acts as the source of truth for translation progress. The Python process does not need to maintain a separate progress file.
python translate_mcqs.py
.
.
Translating MCQ: 000a3922-f4b6-47dd-a686-c5ead8337bb0
Raw model response:
{
"question": "рдмреЗрддрд▓рд╛ рд░рд╛рд╖реНрдЯреНрд░реАрдп рдЙрджреНрдпрд╛рди рдХрд╣рд╛рдБ рд╕реНрдерд┐рдд рд╣реИ?",
"optionA": "рдЙрддреНрддрд░ рдкреНрд░рджреЗрд╢",
"optionB": "рдЭрд╛рд░рдЦрдВрдб",
"optionC": "рдордзреНрдп рдкреНрд░рджреЗрд╢",
"optionD": "рдУрдбрд┐рд╢рд╛",
"explanation": "\"рдмреЗрддрд▓рд╛ (рдкрд╛рд▓рдореБ) рд░рд╛рд╖реНрдЯреНрд░реАрдп рдЙрджреНрдпрд╛рди рдЭрд╛рд░рдЦрдВрдб рдХреЗ рджреЗрд░рд╣рд╛рддрд░ рдЬрд┐рд▓реЗ рдореЗрдВ рд╕реНрдерд┐рдд рд╣реИ, рдЬреЛ рдкрд╣рд▓реЗ рдмрд┐рд╣рд╛рд░ рдХрд╛ рд╣рд┐рд╕реНрд╕рд╛ рдерд╛ред\""
}
.
.
4. Conclusion
This approach provided a practical way to translate approximately 12,000 English MCQs into Hindi without relying on paid LLM APIs for every request.The evaluation showed that TranslateGemma 4B was better suited to this translation workload than Qwen3 4B.
Running it locally through Ollama made the translation process inexpensive, while Python and PostgreSQL provided a simple way to process untranslated records and store the Hindi content alongside the original English content.
The translated flag also made the process restartable, allowing the application to continue from where it stopped without retranslating completed records.