The language gap starts at the keyboard.
A person might type “samsad” while looking for news about parliament. Matching that spelling directly against English articles misses the intended topic. Variations in transliteration and Nepali–English code-mixing add another layer of ambiguity.
NLTTA separates that problem into two tasks: understand the query, then retrieve relevant documents. Translation changes the language of the text; transliteration changes its script. The training corpus includes both Devanagari and Romanized Nepali alongside English.
The project’s practical demonstration is Notify Nepal, a news application that exposes this pipeline through a familiar search interface.
A language model, with somewhere useful to go.
- Enter a queryRomanized or Nepali input
- TranslateFine-tuned mBART
- RankBM25 over news content
- ReadRelevant article results
The translation model is served through a Django REST API. Its English output becomes the query for a custom BM25 ranker; the backend returns matching feed items to a Next.js interface backed by MySQL.
This division makes failures easier to examine. An incorrect translation is a language-model issue. A reasonable translation with poor results points to the news corpus or ranking. BM25 keeps the retrieval step inspectable, using term frequency, document frequency, and document length rather than an opaque generated answer.
The report specifies BM25 parameters of k1 = 1.5 and b = 0.8. The public search endpoint requests the five highest-ranked documents.
The dataset is part of the engineering.
The team combined Wikipedia-derived material with open-source Nepali–English parallel data, then used rule-based Indic transliteration to produce the Romanized column. The report describes approximately 40,000 aligned samples across English, Devanagari Nepali, and Romanized Nepali.
The corpus was split before evaluation: 70% for training, 20% for validation, and 10% for testing. Fine-tuning ran in checkpointed sessions to work within the available compute.
What the loss trace tells us
One documented training session moves from a loss of 1.1967 at step 500 to 0.3449 at step 7,000. That records better fit to the training objective; the separate translation evaluations are needed to judge output quality.
Inspect the original training trace in the reportTwo measurements. Two different questions.
| Measurement | Reported result | Evaluation context |
|---|---|---|
| BLEU | 39.84 | Translation overlap on the held-out 10% test split. |
| BERTScore F1 | 0.9613 | Semantic similarity on randomly selected and generated examples described in §5.3.3. |
BLEU checks overlap with reference translations. BERTScore compares contextual meaning. Together they provide complementary evidence, but the two reported scores do not come from an identically described evaluation set.
The report also compares prompted outputs from general-purpose models. I treat those comparisons as exploratory: model versions, prompts, and evaluation configuration matter, so they do not establish a universal ranking or a production-quality guarantee.
Functional tests cover empty input, stopword-only queries, ranking behavior, and no-result responses. Translation quality and retrieval usefulness remain separate things to test.
From an informal query to readable results.


These are recorded project demonstrations. The repository contains the application code; the screenshots do not represent a live translation service running inside this portfolio.
What I would strengthen next.
Rule-generated Romanization gives us useful parallel data, but it cannot capture every spelling habit, dialect, or informal phrase found in real messages. More naturally occurring Romanized queries would make the evaluation closer to actual use.
- Publish a versioned test set, metric configuration, and reproducible evaluation script.
- Measure retrieval relevance separately from translation quality, including queries that should return no result.
- Expand coverage of slang, regional expressions, and mixed-language input.
- Investigate speech input and other information domains after establishing a stronger evaluation baseline.
The most useful lesson was building the whole path from data to model to interface. A language model becomes more valuable when its output leads someone to information they can use.
The hardest decision: understanding how people actually type.
The hardest part for me was bridging Romanized Nepali with Devanagari when people write the same word—or even the same sound—in several different ways. A consistent transliteration rule is useful for preparing data, but everyday typing is much less consistent.
That tension shaped how I think about the project: a working pipeline is only part of the answer. The next improvement I would prioritize is testing naturally occurring spelling variants, so evaluation reflects the people using the system.
The work behind the write-up.
Final-year B.Sc. CSIT project at Bhaktapur Multiple Campus, Tribhuvan University. Prepared by Bibek Rawal, Prabesh Gautam, and Rahul Koju, supervised by Sushant Paudel. The report is dated November 2025.