PAKGOV-RAG
In progress Version 0 pilotA bilingual Urdu-English retrieval-augmented generation system for Pakistani government and legal documents.
- The problem
- Official information such as tax circulars, education policy, identity procedures, company filings and court judgments is spread across many documents and is hard for Urdu readers to search. Answers from general AI tools may sound confident without pointing to the source document.
- Why it matters
- People should be able to ask in the language they actually write, and check the answer against an official source.
- Document areas targeted
- FBR circulars, HEC policy, NADRA procedures, SECP filings and Supreme Court judgments. Source links and document IDs are kept with each document.
- Approach
-
The repository describes four components, shown below in the order a document passes through them.
- OCRTurn scanned documents into text.
- Hybrid retrievalCombine dense and keyword search to find passages.
- Citation groundingTie each answer to the passages it used.
- EvaluationMeasure retrieval and answer quality.
Components are taken from the repository description. Detailed architecture notes will be added as they are written up.
- Technologies in use now
- Python, Git, FAISS and LangChain. I am still learning PyTorch and Hugging Face, so I do not list them as part of the working system.
- Planned evaluation
- Recall@5, Recall@10, MRR, nDCG, answer correctness, groundedness, citation precision and recall, unsupported claims, hallucination, abstention, speed and cost.
- Measured results
- None published yet. I will add numbers here only after the evaluation has actually been run, with the method and dataset described.
- Known limitations
- This is a pilot, not a production service. It gives no legal advice and I make no claim of legal accuracy.
- Next steps
- Upgrade step by step into PAKGOV-RAG-X: retrieval metrics, source grounding, an evaluation set, and a written technical report.