PIPE-RDF builds schema-specific natural-language-to-SPARQL benchmarks for RDF knowledge graphs. It grounds generation in the target graph, balances query categories, validates every SPARQL query by parsing and execution, and writes reproducible run artifacts for benchmark construction and downstream NL-to-SPARQL evaluation.
This branch is the public code branch. Paper drafts, submission packages, and anonymous-review source files are intentionally not kept here.
Natural-language access to RDF knowledge graphs depends on evaluation sets whose queries actually run on the graph being tested. Existing KGQA benchmarks provide useful shared tasks, but their schemas, namespaces, predicates, and query distributions often differ from the graphs used in practice. We introduce PIPE-RDF, an execution-grounded workflow for building schema-specific natural-language/SPARQL benchmarks from a target RDF graph. PIPE-RDF starts with reverse queries that populate deterministic binding banks, uses category-aware retrieval to supply schema-matched examples, and applies controlled LLM generation to produce candidate question-query pairs. Each candidate is accepted only after passing predicate and type checks, deduplication, parsing, execution, answer-shape validation, and non-empty-result checks. Across a compact company-location schema and a 25M-triple LDBC Semantic Publishing Benchmark graph, PIPE-RDF produces 3,600 balanced pairs across nine query categories with no parse, execution, or empty-answer failures in the released artifacts. Cross-model probes, a dual LLM-judge audit, and downstream prompting experiments indicate that the resulting benchmarks are robust, semantically aligned, and useful as schema-matched examples for NL-SPARQL evaluation.
- Reverse-query grounding so generated questions are answerable on the target graph.
- Category-balanced generation across generic, counting, comparative, superlative, ordinal, multi-hop, intersection, difference, and yes/no query types.
- GraphDB/SPARQL validation with strict parse, execution, answer-shape, and category-form checks.
- Binding-bank sampling and batched label/type lookup to avoid repeated expensive SPARQL calls during large runs.
- LLM backends through Ollama or OpenAI-compatible endpoints such as vLLM.
- Local embeddings through
sentence-transformerswithBAAI/bge-m3. - Utilities for schema profiling, benchmark summarization, semantic LLM judging, and downstream utility evaluation.
pipekg/: Core pipeline modules.configs/: Smoke, full-run, ARR-scale, cross-model, and top-up run configurations.scripts/: GraphDB, vLLM, generation, audit, evaluation, and summarization scripts.db_setup.md: GraphDB setup and loading notes.AGENTS.md: Operational instructions for Codex agents working in this repository.
Generated datasets, logs, GraphDB data, local model outputs, and paper/submission workspaces are ignored by Git.
Create an environment and install dependencies:
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txtCopy the example environment file:
cp .env.example .envSet the SPARQL endpoint for your GraphDB repository:
SPARQL_ENDPOINT_URL=http://localhost:7200/repositories/spb_1mFor GraphDB setup and data loading, see db_setup.md.
For Ollama:
LLM_PROVIDER=ollama
OLLAMA_BASE_URL=http://localhost:11434
OLLAMA_CHAT_MODEL=qwen3:4b-instruct
OLLAMA_EMBED_MODEL=bge-m3:latestFor vLLM or another OpenAI-compatible server:
LLM_PROVIDER=openai_compatible
OPENAI_BASE_URL=http://localhost:8000/v1
OPENAI_CHAT_MODEL=Qwen/Qwen3.5-4B
OPENAI_API_KEY=EMPTY
EMBED_PROVIDER=sentence_transformers
LOCAL_EMBED_MODEL=BAAI/bge-m3Smoke-test an OpenAI-compatible endpoint:
python scripts/vllm_smoke.py --base-url http://localhost:8000/v1 --model Qwen/Qwen3.5-4BOn ds-serv6, the helper script manages GraphDB lifecycle and health checks:
bash scripts/ds_serv6_graphdb.sh status
bash scripts/ds_serv6_graphdb.sh health
bash scripts/ds_serv6_graphdb.sh restartFor local development, start GraphDB using the instructions in db_setup.md, then verify the configured endpoint:
python scripts/verify_endpoint.pyRun a config-level preflight before launching an experiment:
python scripts/preflight_check.py --config configs/smoke_test.yamlSmoke run:
python scripts/run_pipeline_ollama.py --config configs/smoke_test.yamlARR-scale Schema C run:
python scripts/run_pipeline_ollama.py --config configs/arr_schema_c_200.yamlARR-scale SPB run:
python scripts/run_pipeline_ollama.py --config configs/arr_spb_full_200.yamlCross-model probes:
python scripts/run_pipeline_ollama.py --config configs/arr_cross_model_schema_c_50.yaml
python scripts/run_pipeline_ollama.py --config configs/arr_cross_model_spb_full_50.yamlOn ds-serv6, use tmux-backed helper scripts for longer runs:
bash scripts/ds_serv6_run_arr_experiments.shProfile a schema:
python scripts/profile_schema.py --config configs/arr_spb_full_200.yamlSummarize a benchmark artifact:
python scripts/summarize_benchmark_artifact.py --input path/to/phase3.jsonlSample a semantic-audit packet:
python scripts/sample_semantic_audit.py --input path/to/phase3.jsonl --output audit_packet.csvRun dual LLM semantic judges:
python scripts/evaluate_semantic_llm_judges.py \
--input audit_packet.csv \
--output-dir artifacts/llm_semantic_audit/run_name \
--judge openai \
--judge xaiRun downstream utility evaluation:
python scripts/evaluate_downstream_utility.py --help- Keep
.envlocal and never commit API keys or GraphDB credentials. - Keep generated artifacts under ignored directories such as
artifacts/,experiments/, orresults/. - Record the exact config, model, endpoint, GraphDB repository, and run manifest for each reported experiment.
- Keep paper drafts and submission packages outside
main; use dedicated paper/review branches for those.