Abstract
Tamil is an agglutinative and highly inflectional language with rich morphology. Sandhi errors are a class of spelling mistakes in Tamil that occur at word boundaries due to incorrect insertion or omission of hard consonants. This paper presents TamilSandhi, an open-source neuro-symbolic framework for identifying and correcting such errors. It combines rule-based logic for simple Sandhi rules with neural models for one complex rule. The rule-based system implements 12 Vallinam (hard consonant) addition rules and 8 deletion rules, derived from classical Tamil grammar texts such as Tolkappiyam and Nannool. These rules were packaged into a PyPi library (https://pypi.org/project/tamilsandhi-toolkit/) and validated using a suite of 300 unit tests (235 for addition, 65 for deletion). To address a complex rule not covered by the rule-based logic, a neural sequence-to-sequence approach was used. A supervised corpus of 10,434 manually annotated sentence pairs, each consisting of an incorrect and corrected version, was created. Three multilingual transformer models—mBART, mT5, and NLLB—were fine-tuned. Among them, mBART achieved the best performance, with a BLEU score of 99.9 and exact match accuracy of 97.9%.TamilSandhi is released as an open-source project on GitHub (https://github.com/TamilGeekGirl/TamilSandhiNeuroSymbolicAI/). Due to its modularity, reproducibility, and linguistic validity, TamilSandhi constitutes a significant contribution to NLP research in a widely used but low-resource language.
| Original language | English |
|---|---|
| Title of host publication | Artificial Intelligence XLII - 45th SGAI International Conference on Artificial Intelligence, AI 2025, Proceedings |
| Editors | Max Bramer, Frederic Stahl |
| Publisher | Springer Science and Business Media Deutschland GmbH |
| Pages | 243-256 |
| Number of pages | 14 |
| ISBN (Electronic) | 9783032114020 |
| ISBN (Print) | 9783032114013 |
| DOIs | |
| Publication status | Published - 2026 |
| Event | 45th SGAI International Conference on Artificial Intelligence, AI 2025 - Cambridge, United Kingdom Duration: 16 Dec 2025 → 18 Dec 2025 |
Publication series
| Name | Lecture Notes in Computer Science |
|---|---|
| Volume | 16301 LNAI |
| ISSN (Print) | 0302-9743 |
| ISSN (Electronic) | 1611-3349 |
Conference
| Conference | 45th SGAI International Conference on Artificial Intelligence, AI 2025 |
|---|---|
| Country/Territory | United Kingdom |
| City | Cambridge |
| Period | 16/12/25 → 18/12/25 |
Keywords
- Neuro-symbolic AI
- NLG
- Sandhi errors
- Tamil
ASJC Scopus subject areas
- Theoretical Computer Science
- General Computer Science
Fingerprint
Dive into the research topics of 'TamilSandhi: A Neuro-Symbolic AI Toolkit for Correcting Sandhi Errors in Tamil'. Together they form a unique fingerprint.Student theses
-
A Neuro-Symbolic Approach to Address Sandhi and Context-Sensitive Errors in Tamil: Evaluating Deep Learning Models for a Highly Inflectional Language
Vasuki Murugesan, Y. V. M. (Author), Waller, A. (Supervisor) & Visser, J. (Supervisor), 2026Student thesis: Doctoral Thesis › Doctor of Philosophy
File
Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver