AI Engineer · NLP Researcher

Hi, I'm Sami

AI engineer and early-career NLP researcher working on LLMs for low-resource, dialectal languages.

Open to research collaborations and graduate positions in low-resource and dialectal NLP, LLM evaluation, and agentic AI and agent security.

🌿
Try:
2First-author papers
9Projects shipped & in progress
K. M. Jubair Sami

About

Research and engineering

I'm an AI engineer at Link3 Technologies in Dhaka, Bangladesh's largest ISP, where I build LLM agents and agent harnesses, and the data pipelines and production systems around them.

My research looks at how LLMs treat low-resource languages and their dialects, starting with Bengali. My first-author work on Bengali dialectal bias appears at EMNLP 2026 (main conference) and the BLP-2025 workshop. It grew out of my undergraduate thesis at BRAC University, which won the Best Thesis Award. I graduated in Computer Science with a 3.95 CGPA.

Outside work I build MappedTech, a laptop comparison platform with an AI agent that cites its sources. I've made tech videos on YouTube since 2017.

Research interests

Primary

Low-resource and dialectal NLP

  • How LLMs behave on low-resource languages and their dialects, Bengali first: evaluation, bias, fairness.
  • Cost-efficient LLMs for low-resource settings: tokenizer inefficiency (the "token tax") and making LLMs cheap enough to be useful where they're needed most.
  • Mitigating dialectal performance gaps, not just measuring them.
  • Evaluation without references: when to trust an LLM judge, human-in-the-loop evaluation, metrics for unstandardized orthography.

Secondary

Agentic AI and agent security

  • Agent harnesses and the engineering of reliable, tool-using agents.
  • Security of agents that hold sensitive data: prompt injection, misuse, trustworthy and privacy-preserving agents.
Low-resource NLPDialectsBengali / BanglaLLM evaluationLLM-as-judgeFairness & biasTokenizationEfficient LLMsRAGMachine translationAgentic AIAgent security

Skills

Research

LLM evaluation & benchmark designLLM-as-judge with human validation (RLAIF-style)RAG for low-resource machine translationLarge-scale evaluation (68k+ runs, 19 LLMs)Human-in-the-loop annotation

AI engineering

LLM agents (LangGraph)Voice agents (Verbex)RAG & document QAEmbedding searchNon-generative decision models (Jev)Cost vs. accuracy model selection

ML & data

PythonPyTorchTensorFlowNumPyPandasScikit-learnFAISSPineconeChromaDBWhisper ASRData pipelinesChurn prediction

Software

SQLJavaScriptNext.jsFastAPIFlaskREST APIsPostgreSQLMSSQLRBAC

DevOps & observability

DockerCI/CD (Azure DevOps, GitHub)Bare-metal UbuntuPrometheusGrafanaGit

Product & communication

Product owner on 6 internal productsPublic speaking & video since 2017Teaching (CSE221 Algorithms)

Research

Publications

First-author work on how LLMs handle low-resource Bengali dialects. Google Scholar ↗

EMNLP 2026 · Main ConferenceAcceptedBudapest, Hungary · Oct 24–29, 2026

Benchmarking Bengali Dialectal Bias: A Multi-Stage Framework Integrating RAG-Based Translation and Human-Augmented RLAIF

K. M. Jubair Sami, Dipto Sumit, Ariyan Hossain, Farig Sadeque

  • RAG-based translation of standard-Bengali questions into 9 dialects: 4,000 gold-labeled question sets.
  • 19 LLMs benchmarked in 68,000+ evaluation runs; LLM-as-judge validated by 37 native-speaker annotators.
  • Chittagong dialect scores 5.44/10 vs. Tangail 7.68/10; larger models don't reliably reduce the gap.
  • New metric, Critical Bias Sensitivity (CBS), for safety-critical use.
Abstract

Large language models (LLMs) frequently exhibit performance biases against regional dialects of low-resource languages. However, frameworks to quantify these disparities remain scarce. We propose a two-phase framework to evaluate dialectal bias, operationalized as comprehension degradation relative to standard Bengali, in LLM question-answering across nine Bengali dialects. First, we translate and gold-label standard Bengali questions into dialectal variants adopting a retrieval-augmented generation (RAG) pipeline to prepare 4,000 question sets. Since traditional translation quality evaluation metrics fail on unstandardized dialects, we evaluate fidelity using an LLM-as-a-judge, which human correlation confirms outperforms legacy metrics. Second, we benchmark 19 LLMs across these gold-labeled sets, running 68,395 RLAIF evaluations validated through multi-judge agreement and human fallback. Our findings reveal severe performance drops linked to linguistic divergence. For instance, responses to the highly divergent Chittagong dialect score 5.44/10, compared to 7.68/10 for Tangail. Furthermore, increased model scale does not consistently mitigate this bias. We contribute a validated translation quality evaluation method, a rigorous benchmark dataset, and a Critical Bias Sensitivity (CBS) metric for safety-critical applications.

Extended version of my undergraduate thesis (Best Thesis Award, BRAC University, Fall 2025).

Low-resource NLPDialectsLLM evaluationBiasLLM-as-judgeRAG
BLP-2025 @ IJCNLP-AACL · WorkshopPublishedMumbai, India · December 2025

A Comparative Analysis of Retrieval-Augmented Generation Techniques for Bengali Standard-to-Dialect Machine Translation Using LLMs

K. M. Jubair Sami, Dipto Sumit, Ariyan Hossain, Farig Sadeque

  • Standard Bengali → 6 regional dialects with very little data and no fine-tuning.
  • Structured sentence-pair retrieval beats transcript context: Chittagong WER falls from 76% to 55%.
  • Retrieval design matters more than model size: Llama-3.1-8B beats much larger models.
Abstract

Translating from a standard language to its regional dialects is a significant NLP challenge due to scarce data and linguistic variation, a problem prominent in the Bengali language. This paper proposes and compares two novel RAG pipelines for standard-to-dialectal Bengali translation. The first, a Transcript-Based Pipeline, uses large dialect sentence contexts from audio transcripts. The second, a more effective Standardized Sentence-Pairs Pipeline, utilizes structured local_dialect:standard_bengali sentence pairs. We evaluated both pipelines across six Bengali dialects and multiple LLMs using BLEU, ChrF, WER, and BERTScore. Our findings show that the sentence-pair pipeline consistently outperforms the transcript-based one, reducing Word Error Rate (WER) from 76% to 55% for the Chittagong dialect. Critically, this RAG approach enables smaller models (e.g., Llama-3.1-8B) to outperform much larger models (e.g., GPT-OSS-120B), demonstrating that a well-designed retrieval strategy can be more crucial than model size. This work contributes an effective, fine-tuning-free solution for low-resource dialect translation, offering a practical blueprint for preserving linguistic diversity.

Written in my third year of undergrad. Its translation pipeline powers Phase 1 of the EMNLP 2026 paper.

Low-resource NLPDialectsMachine translationRAG

Experience

Where I've worked

  1. AI Engineer

    Feb 2026 – Present

    Link3 Technologies Ltd. · Bangladesh's largest ISP

    Joined as a QA engineer, moved to AI within about two weeks, and became a full-time AI engineer in June 2026. Product owner on 6 internal products for the CeX, HR and CLM departments.

    • AI Engineer (Full-time) · Jun 2026 – Present
    • AI Engineer (Intern) · Feb 2026 – May 2026
    • QA Engineer · Feb 2026
  2. Educational Content Creator

    Dec 2023 – Present

    MappedAcademy

    Educational content on data structures and algorithms for local students.

  3. Tech Content Creator

    Jan 2017 – Present

    YouTube · MappedTech

    Tech reviews, and later laptop buying guides for Bangladesh. The channel started as a way to get over stage fright; the brand continues as MappedTech.

  4. Student Tutor, CSE221 Algorithms

    Feb 2025 – May 2025

    BRAC University

    Undergraduate teaching assistant for the Algorithms course for one semester.

Contact

Get in touch

Open to research collaborations and graduate positions in low-resource and dialectal NLP, LLM evaluation, and agentic AI and agent security.

Based in Dhaka, Bangladesh

🌿

Mimi

Online

Hi! I'm Mimi 🌿, Sami's AI assistant. Ask me about his research, projects, experience or background.

  • "What's Sami's latest paper about?"
  • "What is ChurnSync?"
  • "Where does Sami work now?"