Trustworthy AI for Urdu and Pakistani public information.

I am Salik Hussain, a first-year BS Computer Science student at Iqra University Islamabad. I am building PAKGOV-RAG, a bilingual Urdu-English retrieval-augmented generation (RAG) system for Pakistani government documents, and studying whether AI answers equally well in Urdu script, Roman Urdu and mixed Urdu-English text.

Based in Rawalpindi, Pakistan. Available for freelance Python, web and AI projects.

The research problem in one example
Urdu script
پاسپورٹ کی فیس کتنی ہے؟
Roman Urdu
Passport ki fee kitni hai?
Urdu-English mixed
پاسپورٹ کی fee کتنی ہے؟

Illustrative inputs only. This is not model output and no results are claimed. The open question is whether a system finds and cites the same correct source for all three.

Peer-reviewed publications
None yet
Core papers planned
3
Public GitHub repositories
29
GitHub contributions, past year
719

GitHub figures were read from my profile in October 2026. Everything on this page is finished, in progress, or marked Planned.

Research focus

One question guides my work: can AI give correct, trusted answers in Urdu, Roman Urdu and Urdu-English mixed text?

Status labels used on this page: Active In progress Planned

  • Urdu-RomanX

    Planned

    Do models understand Urdu script, Roman Urdu and Urdu-English text equally well?

  • UrduQA-Reason

    Planned

    Can AI answer with evidence and multi-step reasoning, and say so when no answer exists?

  • UrduCodeSwitch-Bench

    Planned

    How reliable and safe is AI across different forms of the same language?

  • UrduIE-KG

    Planned

    Can we extract people, places, laws and relations from Urdu text?

  • PAKGOV-RAG-X

    Planned

    Can AI answer Pakistani government questions with real, checkable sources? This is the planned research upgrade of the current PAKGOV-RAG prototype.

Wider interests: Nastaliq text processing, bilingual embeddings, code-switching, and low-resource fine-tuning with models such as mBERT, XLM-RoBERTa and mT5. Stretch work on RAG reliability benchmarks follows only after the core projects.

Projects

One flagship case study, followed by the other repositories pinned on my GitHub profile.

PAKGOV-RAG

In progress Version 0 pilot

A bilingual Urdu-English retrieval-augmented generation system for Pakistani government and legal documents.

The problem
Official information such as tax circulars, education policy, identity procedures, company filings and court judgments is spread across many documents and is hard for Urdu readers to search. Answers from general AI tools may sound confident without pointing to the source document.
Why it matters
People should be able to ask in the language they actually write, and check the answer against an official source.
Document areas targeted
FBR circulars, HEC policy, NADRA procedures, SECP filings and Supreme Court judgments. Source links and document IDs are kept with each document.
Approach
The repository describes four components, shown below in the order a document passes through them.
  1. OCRTurn scanned documents into text.
  2. Hybrid retrievalCombine dense and keyword search to find passages.
  3. Citation groundingTie each answer to the passages it used.
  4. EvaluationMeasure retrieval and answer quality.

Components are taken from the repository description. Detailed architecture notes will be added as they are written up.

Technologies in use now
Python, Git, FAISS and LangChain. I am still learning PyTorch and Hugging Face, so I do not list them as part of the working system.
Planned evaluation
Recall@5, Recall@10, MRR, nDCG, answer correctness, groundedness, citation precision and recall, unsupported claims, hallucination, abstention, speed and cost.
Measured results
None published yet. I will add numbers here only after the evaluation has actually been run, with the method and dataset described.
Known limitations
This is a pilot, not a production service. It gives no legal advice and I make no claim of legal accuracy.
Next steps
Upgrade step by step into PAKGOV-RAG-X: retrieval metrics, source grounding, an evaluation set, and a written technical report.

Other repositories

  • Daily Python practice, programming fundamentals, algorithms and projects, documenting my progress from beginner to proficient.

    • Python
  • A reproducible, source-grounded engine that models program-specific admission requirements from official university sources and validates applicant inputs.

    • Python
  • Bilingual text analyzer for English and Roman Urdu text, with word-frequency analysis, UTF-8 support and CSV output.

    • Python
  • Command-line tool for fast keyword search across Pakistani government documents, with contextual results, search history, error handling and modular file processing.

    • Python
  • Reproducing landmark machine learning papers from scratch, including Attention Is All You Need and BERT.

See all repositories on GitHub

Research and publications

I have no peer-reviewed publications yet. A paper is an outcome, not a guarantee, and I will not submit work only to reach a number. All papers below are Planned.

Planned core papers
Working titleBuilds onTarget timingStatus
Urdu-RomanX: Robust Representation Learning Across Urdu, Roman Urdu and Urdu-English TextUrdu-RomanXDraft in year 2, submit in year 3Planned
UrduCodeSwitch-Bench: Reliability, Factuality and Safety Evaluation Across Urdu, Roman Urdu and Code-Switched TextUrduCodeSwitch-BenchYears 3 to 4Planned
PAKGOV-RAG-X: Evidence-Grounded Bilingual RAG for Pakistani Government and Public InformationPAKGOV-RAG-XYear 4Planned

Venues are chosen only after the real contribution is clear. My research identity is on ORCID (0009-0002-2163-1728).

Technical skills

Grouped by how far along I am, so you can see what I can use today and what I am still building.

Working knowledge

Used in projects or certified.

  • Python
  • C
  • SQL
  • HTML and CSS
  • JavaScript (basic)
  • Git and GitHub
  • Linux
  • FAISS
  • LangChain
  • REST APIs

Learning now

Studied or in early use.

  • C++
  • PyTorch
  • Hugging Face
  • NumPy and pandas
  • Calculus
  • Linear algebra
  • Probability

Planned next

Scheduled in my study plan.

  • scikit-learn
  • Transformers and embeddings
  • RAG evaluation
  • FastAPI
  • Docker
  • Experiment tracking

Background

I am from Kurram District and study in Islamabad. I started in pre-medical (F.Sc. PCB), then moved to computer science because I saw how hard it is for Urdu readers to use government information.

I build practical tools, learn in public, and aim for a fully funded MS in AI abroad. Freelance work lets me practise on real problems while I study. Targets on this page are goals, not achievements.

Education

  • BS Computer Science, Iqra University Islamabad
    Started October 2026. Fall 2026 courses: Programming Fundamentals (C++), Application of ICT, Functional English, Applied Physics, Pre-Calculus I.
  • F.Sc. Pre-Medical
    HSSC 71.18 percent, completed August 2026.
  • SSC
    83.27 percent.

Study plan

A plan, not a record. Dates move if grades or opportunities change.

  1. Year 1, 2026 to 2027Strong grades, solid C++ and calculus, daily Git habit, first research reading, and a review of the PAKGOV-RAG version 0 prototype.
  2. Year 2, 2027 to 2028Data structures, algorithms and machine learning basics. Finish the first Urdu NLP research project and draft the first paper.
  3. Year 3, 2028 to 2029AI core courses, an Urdu code-switching benchmark, research internship applications and English test preparation.
  4. Year 4, 2029 to 2030PAKGOV-RAG-X as the final-year project and MS applications for fully funded programs. No one can promise admission.

Credentials

24 courses and assessments are listed. Certificates show learning. They are not research results. Where a public verification link exists, it is linked.

University and platform courses, issued August and September 2026
CourseProviderIssuedVerification
CS50x: Introduction to Computer ScienceHarvardAug 2026Link not added yet
CS50 Artificial Intelligence with PythonHarvardAug 2026Link not added yet
CS50 Databases with SQLHarvardAug 2026Link not added yet
CS50 Introduction to CybersecurityHarvardSep 2026Link not added yet
Artificial Intelligence Fundamentals
ID bfe0dc1d-e245-444b-b046-c4d620b16bd2
IBM SkillsBuildSep 2026On LinkedIn
Getting Started with Generative AI
ID PWID-B1036800
IBM SkillsBuildSep 2026On LinkedIn
Machine Learning Crash Course (numerical data module)GoogleSep 2026On LinkedIn
Introduction to Programming in CThe Open UniversitySep 2026On LinkedIn
Elements of AI (2 ECTS)
ID vh5jd0sod38
University of HelsinkiAug 2026On LinkedIn
Claude Platform 101AnthropicAug 2026On LinkedIn
Skill assessments and other credentials
CredentialProviderVerification
CS50's Introduction to Programming with PythonHarvardVerify certificate
Problem Solving (Intermediate)HackerRankVerify certificate
Problem Solving (Basic)HackerRankVerify certificate
Python (Basic)HackerRankVerify certificate
SQL (Advanced)HackerRankVerify certificate
SQL (Intermediate)HackerRankVerify certificate
SQL (Basic)HackerRankVerify certificate
REST API (Intermediate)HackerRankVerify certificate
JavaScript (Basic)HackerRankVerify certificate
Software Engineer role certificateHackerRankVerify certificate
Scientific Computing with PythonfreeCodeCampNo public link yet
Introduction to LLMsSoloLearnNo public link yet
Machine Learning for BeginnersSoloLearnNo public link yet
HP LIFE AmbassadorHP LIFENo public link yet

Practice profiles: LeetCode, HackerRank and Kaggle.

Freelance work

Fixed-scope work with a written quote before I start. You approve each stage before the next begins. Studies come first, so I take only projects I can deliver well.

  • Python automation

    Scripts that clean files, collect public data, send reports and save manual work.

    • Python
    • Pandas
    • APIs
  • Websites

    Fast, responsive business and portfolio sites with clear copy and working contact forms.

    • HTML
    • CSS
    • JavaScript
  • AI chatbot and document search

    Ask questions over your PDFs and documents in English or Urdu, with the source shown for each answer.

    • RAG
    • LangChain
    • FAISS
  • Data cleaning and analysis

    Turn messy spreadsheets and CSV files into clean data and clear charts.

    • SQL
    • Pandas
  • Urdu and English text work

    Search, classification and processing for Urdu text, including Nastaliq documents.

    • NLP

How we work

  1. Brief

    Tell me what you need. I reply with questions and a fixed quote within 24 hours.

  2. Build

    I share short progress updates so you see work early, not only at the end.

  3. Review

    You test it and request changes. Revisions are included in the agreed scope.

  4. Deliver

    Clean code, a short guide and a handover. I stay available for fixes afterwards.

Questions clients ask

Are you experienced enough?

I am a first-year BS CS student with certificates in Python, SQL, REST APIs and JavaScript, and an AI project on GitHub. I tell you early if something is outside my skills.

How fast can you start?

Usually within two to three days of agreeing the scope. Small projects can finish within a week.

How do we communicate and pay?

Email, WhatsApp or the platform you prefer. Payment goes through Upwork, Fiverr or an agreed method, with milestones for larger work.

Do you work with Urdu content?

Yes. Urdu and English text processing, Urdu search and Urdu-English document question answering are my focus.

Contact

For research advice, supervision enquiries or freelance work, email me. I reply within 24 hours.

Email me

[email protected]

Profiles

Send an enquiry

This opens your email app with the message filled in. Nothing is sent from this page until you press send in your email app.