# Pralekha: Cross-Lingual Document Alignment for Indic Languages
arXiv HuggingFace GitHub License: CC BY 4.0
# Overview **PRALEKHA** is a large-scale document-level benchmark for Cross-Lingual Document Alignment (CLDA) evaluation, comprising over 3 million aligned document pairs across 11 Indic languages and English, with 1.5 million being English-Indic pairs. We propose a comprehensive evaluation framework that introduces **Document Alignment Coefficient (DAC)**, a novel metric specifically designed for fine-grained document alignment. Unlike existing approaches that use pooled document-level embeddings, DAC aligns smaller chunks within documents and computes similarity based on the ratio of aligned chunks. This method outperforms baseline pooling approaches, particularly in noisy scenarios, yielding 15–20% precision and 5–10% F1 score improvements. # Usage ### 1. Setup Follow these steps to set up the environment and get started with the pipeline: #### i. Clone the Repository Clone this repository to your local system: ```python git clone https://github.com/AI4Bharat/Pralekha.git cd Pralekha ``` #### ii. Set Up a Conda Environment Create and activate a new Conda environment for this project: ```python conda create -n pralekha python=3.9 -y conda activate pralekha ``` #### iii. Install Dependencies Install the required Python packages: ```python pip install -r requirements.txt ``` ### 2. Input Directory Structure The pipeline expects a directory structure in the following format: - A **main directory** containing language subdirectories named using their **3-letter ISO codes** (e.g., `eng` for English, `hin` for Hindi, `tam` for Tamil, etc.) - Each language subdirectory will contain `.txt` documents named in the format `{doc_id}.txt`, where `doc_id` serves as the unique identifier for each document. Below is an example of the expected directory structure: ```plaintext data/ ├── eng/ │ ├── tech-innovations-2023.txt │ ├── sports-highlights-day5.txt │ ├── press-release-456.txt │ ├── ... ├── hin/ │ ├── daily-briefing-april.txt │ ├── market-trends-yearend.txt │ ├── इंडिया-न्यूज़123.txt │ ├── ... ├── tam/ │ ├── kollywood-review-movie5.txt │ ├── 2023-pilgrimage-guide.txt │ ├── கடலோர-மாநில-செய்தி.txt │ ├── ... ... ``` ### 3. Split Documents into Granular Shards To process documents into granular shards, use the `doc2granular-shards.sh` script. This script splits documents into chunks of varying granularities: - **G = 1** → Sentence-level - **G = 2, 4, 8** → Chunk-level (2, 4, or 8 sentences per chunk) Run the script: ```bash bash doc2granular-shards.sh ``` ### 4. Create Embeddings Generate embeddings for your dataset using one of the two supported models: `LaBSE` or `SONAR`. ```bash bash create_embeddings.sh ``` Choose the desired model by editing the script as needed. Both models can be run sequentially or independently by enabling/disabling the respective sections. ### 5. Run the Pipeline The final step is to execute the pipeline based on your chosen method: For `baseline` approaches: ```bash bash run_baseline_pipeline.sh ``` For the proposed `DAC` approach: ```bash bash run_dac_pipeline.sh ``` Each pipeline comes with a variety of configurable parameters, allowing you to tailor the process to your specific requirements. Please review and edit the scripts as needed before running to ensure they align with your desired configurations. # License This dataset is released under the [**CC BY 4.0**](https://creativecommons.org/licenses/by/4.0/) license. # Citation If you use Pralekha in your work, please cite us: ``` @inproceedings{suryanarayanan-etal-2025-pralekha, title = "{PRALEKHA}: Cross-Lingual Document Alignment for {I}ndic Languages", author = "Suryanarayanan, Sanjay and Song, Haiyue and Khan, Mohammed Safi Ur Rahman and Kunchukuttan, Anoop and Dabre, Raj", booktitle = "Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics", month = dec, year = "2025", address = "Mumbai, India", publisher = "The Asian Federation of Natural Language Processing and The Association for Computational Linguistics", url = "https://aclanthology.org/2025.ijcnlp-long.37/", pages = "662--676" } ``` # Contact For any questions or feedback, please contact: - Sanjay Suryanarayanan ([sanj.ai@outlook.com](mailto:sanj.ai@outlook.com)) - Haiyue Song ([haiyue.song@nict.go.jp](mailto:haiyue.song@nict.go.jp)) - Raj Dabre ([raj.dabre@cse.iitm.ac.in](mailto:raj.dabre@cse.iitm.ac.in)) Please get in touch with us for any copyright concerns.