Hi, I'm Sami
AI engineer and early-career NLP researcher working on LLMs for low-resource, dialectal languages.
Open to research collaborations and graduate positions in low-resource and dialectal NLP, LLM evaluation, and agentic AI and agent security.

About
Research and engineering
I'm an AI engineer at Link3 Technologies in Dhaka, Bangladesh's largest ISP, where I build LLM agents and agent harnesses, and the data pipelines and production systems around them.
My research looks at how LLMs treat low-resource languages and their dialects, starting with Bengali. My first-author work on Bengali dialectal bias appears at EMNLP 2026 (main conference) and the BLP-2025 workshop. It grew out of my undergraduate thesis at BRAC University, which won the Best Thesis Award. I graduated in Computer Science with a 3.95 CGPA.
Outside work I build MappedTech, a laptop comparison platform with an AI agent that cites its sources. I've made tech videos on YouTube since 2017.
Research interests
Primary
Low-resource and dialectal NLP
- How LLMs behave on low-resource languages and their dialects, Bengali first: evaluation, bias, fairness.
- Cost-efficient LLMs for low-resource settings: tokenizer inefficiency (the "token tax") and making LLMs cheap enough to be useful where they're needed most.
- Mitigating dialectal performance gaps, not just measuring them.
- Evaluation without references: when to trust an LLM judge, human-in-the-loop evaluation, metrics for unstandardized orthography.
Secondary
Agentic AI and agent security
- Agent harnesses and the engineering of reliable, tool-using agents.
- Security of agents that hold sensitive data: prompt injection, misuse, trustworthy and privacy-preserving agents.
Skills
Research
AI engineering
ML & data
Software
DevOps & observability
Product & communication
Research
Publications
First-author work on how LLMs handle low-resource Bengali dialects. Google Scholar ↗
Benchmarking Bengali Dialectal Bias: A Multi-Stage Framework Integrating RAG-Based Translation and Human-Augmented RLAIF
- RAG-based translation of standard-Bengali questions into 9 dialects: 4,000 gold-labeled question sets.
- 19 LLMs benchmarked in 68,000+ evaluation runs; LLM-as-judge validated by 37 native-speaker annotators.
- Chittagong dialect scores 5.44/10 vs. Tangail 7.68/10; larger models don't reliably reduce the gap.
- New metric, Critical Bias Sensitivity (CBS), for safety-critical use.
Abstract
Large language models (LLMs) frequently exhibit performance biases against regional dialects of low-resource languages. However, frameworks to quantify these disparities remain scarce. We propose a two-phase framework to evaluate dialectal bias, operationalized as comprehension degradation relative to standard Bengali, in LLM question-answering across nine Bengali dialects. First, we translate and gold-label standard Bengali questions into dialectal variants adopting a retrieval-augmented generation (RAG) pipeline to prepare 4,000 question sets. Since traditional translation quality evaluation metrics fail on unstandardized dialects, we evaluate fidelity using an LLM-as-a-judge, which human correlation confirms outperforms legacy metrics. Second, we benchmark 19 LLMs across these gold-labeled sets, running 68,395 RLAIF evaluations validated through multi-judge agreement and human fallback. Our findings reveal severe performance drops linked to linguistic divergence. For instance, responses to the highly divergent Chittagong dialect score 5.44/10, compared to 7.68/10 for Tangail. Furthermore, increased model scale does not consistently mitigate this bias. We contribute a validated translation quality evaluation method, a rigorous benchmark dataset, and a Critical Bias Sensitivity (CBS) metric for safety-critical applications.
Extended version of my undergraduate thesis (Best Thesis Award, BRAC University, Fall 2025).
A Comparative Analysis of Retrieval-Augmented Generation Techniques for Bengali Standard-to-Dialect Machine Translation Using LLMs
- Standard Bengali → 6 regional dialects with very little data and no fine-tuning.
- Structured sentence-pair retrieval beats transcript context: Chittagong WER falls from 76% to 55%.
- Retrieval design matters more than model size: Llama-3.1-8B beats much larger models.
Abstract
Translating from a standard language to its regional dialects is a significant NLP challenge due to scarce data and linguistic variation, a problem prominent in the Bengali language. This paper proposes and compares two novel RAG pipelines for standard-to-dialectal Bengali translation. The first, a Transcript-Based Pipeline, uses large dialect sentence contexts from audio transcripts. The second, a more effective Standardized Sentence-Pairs Pipeline, utilizes structured local_dialect:standard_bengali sentence pairs. We evaluated both pipelines across six Bengali dialects and multiple LLMs using BLEU, ChrF, WER, and BERTScore. Our findings show that the sentence-pair pipeline consistently outperforms the transcript-based one, reducing Word Error Rate (WER) from 76% to 55% for the Chittagong dialect. Critically, this RAG approach enables smaller models (e.g., Llama-3.1-8B) to outperform much larger models (e.g., GPT-OSS-120B), demonstrating that a well-designed retrieval strategy can be more crucial than model size. This work contributes an effective, fine-tuning-free solution for low-resource dialect translation, offering a practical blueprint for preserving linguistic diversity.
Written in my third year of undergrad. Its translation pipeline powers Phase 1 of the EMNLP 2026 paper.
Experience
Where I've worked
AI Engineer
Feb 2026 – PresentLink3 Technologies Ltd. · Bangladesh's largest ISP
Joined as a QA engineer, moved to AI within about two weeks, and became a full-time AI engineer in June 2026. Product owner on 6 internal products for the CeX, HR and CLM departments.
- AI Engineer (Full-time) · Jun 2026 – Present
- AI Engineer (Intern) · Feb 2026 – May 2026
- QA Engineer · Feb 2026
Educational Content Creator
Dec 2023 – PresentMappedAcademy
Educational content on data structures and algorithms for local students.
Tech Content Creator
Jan 2017 – PresentYouTube · MappedTech
Tech reviews, and later laptop buying guides for Bangladesh. The channel started as a way to get over stage fright; the brand continues as MappedTech.
Student Tutor, CSE221 Algorithms
Feb 2025 – May 2025BRAC University
Undergraduate teaching assistant for the Algorithms course for one semester.
Projects
Selected work
Production AI and data systems at Link3, and MappedTech, my own product. Open any project for the full write-up.
ChurnSync
Churn-intelligence platform over incident data for 335k+ customers.
- 335k+
- customers covered by the pipelines
- 235k+
- active customers
- 1
- engineer (me)
Intelligent Ticketing System
Ticket creation for call-center agents, with the category predicted automatically.
- 2 weeks
- to ship v1 to production
- v2
- Jev integration in progress
MappedTech & Mappie
Laptop comparison and pricing for Bangladesh, with an AI agent that cites its sources.
- 6 h
- price update interval
- ৳40K–150K
- curated budget lists
Contact
Get in touch
Open to research collaborations and graduate positions in low-resource and dialectal NLP, LLM evaluation, and agentic AI and agent security.
- Google Scholar
- GitHub@jubairsami01
- LinkedIn/in/jubairsami
- YouTube · MappedTech@MappedTech
- YouTube · MappedAcademy@MappedAcademy
Based in Dhaka, Bangladesh