A lightweight Retrieval-Augmented Generation (RAG) system with a modern web interface, built in Go.
- R³ Governed Retrieval: Ranked, Responsible, Retrieval with policy-driven scoring and citations
- Semantic Search: Store and search documents using vector embeddings
- RAG Chat: Ask questions and get answers based on your knowledge base
- Multiple Data Sources:
- Wikipedia articles
- Web scraping
- Text input
- File upload (.txt, .md, .csv, .json, .xml, .html, .log)
- Folder import (recursive)
- OpenAI-Compatible API: Works with any OpenAI-compatible LLM backend (LM Studio, Ollama, etc.)
- Custom APIs: Add external API integrations
- Personas: Configure different conversation styles with pre-prompts
- Persistent Agent Memory: Save explicit preferences and stable context for future chats; entries are opt-in and never inferred from transcripts
- Themes: Multiple built-in themes (Dark, Light, Nord, Solarized, Monokai, Dracula)
- Code Execution: Optional support for nanoGo (interpreted Go) execution
- Embedded Frontend: No separate build required - all assets embedded in the binary
# Clone the repository
git clone https://github.com/SimonWaldherr/tinyRAG.git
cd tinyRAG
# Build the application
make build
# Run
./bin/tinyrag -web -addr :8080# Format, vet, and run
make dev
# Or build only
make build
# Run all checks
make check./bin/tinyrag -web -addr :8080Options:
-web: Enable web interface (default: true)-addr: Server address (default: :8080)-db: Database file path (default: data/tinyrag.gob)-settings: Settings JSON path (default: config/settings.json)-url: LLM API base URL (default: http://localhost:1234)-inference-api: inference wire protocol (auto,openai, orollama)-chat-model: Chat model name (first run only)-embed-model: Embedding model name (first run only)-lang: Language code (default: de)-chunk-size: Chunk size for text splitting (default: 800)-k: Number of chunks to retrieve for RAG (default: 5)
The application stores local configuration in config/settings.json, created on
first run and ignored by Git. Start from the tracked
config/settings.example.json if you want to
provision it manually; do not commit API keys or other credentials. You can
also modify the local file through the web interface settings panel.
Example configuration:
{
"version": 3,
"base_url": "http://localhost:1234",
"chat_model": "mistralai/ministral-3-14b-reasoning",
"embed_model": "text-embedding-nomic-embed-text-v1.5",
"lang": "de",
"theme": "monokai",
"chunk_size": 800,
"k": 5,
"custom_apis": [],
"personas": [
{
"id": "persona-default",
"name": "Standard",
"prompt": ""
}
],
"agent_memory_enabled": false,
"agent_memory": [],
"allow_code_exec": false,
"allow_nanogo": false
}Under Settings → Memory, users can save up to 32 explicit, short pieces of durable context, such as language preferences, project conventions, or a default timezone. Memory is disabled by default and is only included in future answers after it has been enabled.
tinyRAG never creates memory entries from conversations, model output, or tool
results. Entries can be added, removed, or disabled in the UI; the protected
API offers the same operations through GET/POST /api/memory and
POST /api/memory/delete.
tinyRAG uses tinySQL v0.49.0. The settings panel exposes the following optional database features:
- Native vector retrieval uses
VEC_SEARCHfor faster candidate lookup. The existing scalar retrieval mode remains available when maximum recall is preferred. - Hybrid retrieval (new) uses tinySQL's
HYBRID_SEARCHto fuse the vector candidate list with a real BM25 full-text pass over chunk content via reciprocal rank fusion, recovering exact identifiers and rare terms that cosine similarity alone can miss. The R³ ranking pipeline still scores every candidate on plain cosine similarity, not the raw fusion score, so downstream thresholds/reranking behave exactly as before. - Configurable vector index (new) picks the ANN strategy
VEC_SEARCH/HYBRID_SEARCHuse:flat(exact, default),ivf, orhnsw. Benchmark before switching away fromflat— corpus size/shape determines the winner. - Startup index warm-up (new) runs tinySQL's
VEC_WARMonce at startup whenvector/hybridretrieval is enabled, so the first real query never pays the one-time index-build cost. A no-op under the defaultscalarmode. - Vector result cache stores only deterministic result IDs, never source
text or embedding vectors. New and migrated installations enable a bounded
cache with 128 entries and a 30-second TTL. Set
tinysql_vector_cache_entriesto0to disable it. - Vector analytics records only query shape and timing. Administrators can
inspect cache and analytics state at
GET /api/debug/vector-cache. - Portable snapshots are available to administrators at
GET /api/debug/database-snapshot. They are consistent tinySQL GOB exports and work independently of the selected storage backend. - Tamper-evident audit logging writes a hash-chained JSONL audit trail.
It can be enabled with
tinysql_audit_enabled; changing it requires a restart. - Encryption at rest for disk, index, and hybrid storage modes reads a
32-byte hexadecimal or Base64 key from
TINYRAG_STORAGE_KEY. The key is never written toconfig/settings.json; enabling encryption requires a restart. - Geodata import accepts GeoJSON, KML, and OpenStreetMap XML through the
Open Data panel when
geo_import_enabledis enabled.
See docs/tinysql-optional-features.md for operational details and constraints.
The universal retrieval design and its staged relation-aware roadmap are in docs/retrieval-architecture.md.
Configured REST, JSON-RPC 2.0, and SQL capabilities are documented in docs/connectors.md. Agent execution admits only the explicitly read-only subset of those capabilities.
tinyRAG talks to any OpenAI-compatible chat/embeddings endpoint, local or
cloud, and can also use the native Ollama /api protocol. The provider switcher in the top toolbar lists common presets
(grouped Local / Cloud) and pre-fills the default base URL for each; picking
one probes the endpoint, lists available models, and lets you apply a chat
model in a couple of clicks. "Custom..." opens Settings for anything not in
the list — any OpenAI-compatible server works even if it isn't listed.
| Provider | Type | Default base URL |
|---|---|---|
| LM Studio | Local | http://localhost:1234 |
| Ollama | Local | http://localhost:11434 |
llama.cpp (llama-server) |
Local | http://localhost:8080 |
| vLLM | Local | http://localhost:8000 |
| text-generation-webui | Local | http://localhost:5000 |
| KoboldCpp | Local | http://localhost:5001 |
| Jan | Local | http://localhost:1337 |
| LocalAI | Local | http://localhost:8080 |
| GopherLLM | Local | http://localhost:8091 (embedded demo default) |
| RustyLLM | Local | http://localhost:8091 (change if your server differs) |
| OpenAI | Cloud | https://api.openai.com |
| Anthropic | Cloud | https://api.anthropic.com |
| Google Gemini | Cloud | https://generativelanguage.googleapis.com |
| Mistral AI | Cloud | https://api.mistral.ai |
| Groq | Cloud | https://api.groq.com/openai |
| DeepSeek | Cloud | https://api.deepseek.com |
| Together AI | Cloud | https://api.together.xyz |
| xAI (Grok) | Cloud | https://api.x.ai |
| Cohere | Cloud | https://api.cohere.ai |
| Perplexity | Cloud | https://api.perplexity.ai |
| OpenRouter | Cloud | https://openrouter.ai/api |
Full setup instructions (install commands, default models, quirks per provider) are in docs/llm-providers.md.
Zero-install demo option: build with
go build -tags demo_llm -o bin/tinyrag ./cmd/tinyrag and run
./bin/tinyrag -demo-llm-model auto to run a tiny pure-Go model
(GopherLLM) in-process — no
LM Studio/Ollama/llama.cpp needed. Demo quality only; see
docs/llm-providers.md.
Quick start with the two most common local runners:
-
LM Studio:
- Download and install LM Studio
- Load a chat model (e.g., Mistral, Llama)
- Load an embedding model (e.g., nomic-embed-text)
- Start the local server (usually runs on port 1234)
-
Ollama:
# Install Ollama curl -fsSL https://ollama.ai/install.sh | sh # Pull models ollama pull llama2 ollama pull nomic-embed-text
-
Configure tinyRAG:
- Open the web interface
- Click the provider switcher in the toolbar and pick your backend, or click the settings (⚙) button → "LLM Backend" tab for manual entry
- Enter your API endpoint (if not using the switcher)
- Click "Test & Load Models"
- Select your chat and embedding models
- Click "Save"
On startup, if the configured endpoint is unreachable, tinyRAG automatically
probes the common local ports above (LM Studio, Ollama, llama.cpp, vLLM,
text-generation-webui, KoboldCpp, Jan) and switches to the first one it finds
— see maybePreferOfflineLLM in llm_discovery.go.
Access the web interface at http://localhost:8080 (or your configured address).
- Chat: Ask questions about your knowledge base
- Search: Perform semantic search on stored chunks
- Data Import: Add documents to your knowledge base
- Wikipedia: Load articles directly
- URL: Scrape web pages
- Text: Paste text content
- Upload: Upload text files
- Folder: Import entire directories
- Chats: View and manage conversation history
- Sources: Browse imported documents
- General: Theme selection and general options
- LLM Backend: Configure API endpoint and models
- Custom APIs: Add external API integrations
- Personas: Create conversation personas with custom prompts
- Memory: Explicitly save, review, enable, and delete durable assistant context
.
├── cmd/tinyrag/ # Minimal executable entry point
├── internal/app/ # Application package and tests
│ ├── web/ # Embedded HTML, CSS, and JavaScript
│ └── examples/ # Embedded static gallery pages
├── config/
│ └── settings.example.json # Safe configuration template
├── data/ # Ignored local DB, chats, uploads, and logs
├── bin/ # Ignored local build outputs
├── docs/ # Operational and architecture documentation
├── go.mod # Go module definition
├── go.sum # Go module checksums
└── Makefile # Build automation
make fmt # Format Go code
make vet # Run go vet
make lint # Run golangci-lint
make tidy # Tidy Go modules
make build # Build bin/tinyrag
make test # Run tests
make check # Run all checks (fmt, vet, lint, test)
make run # Run the application
make dev # Format, vet, and run
make help # Show available targetsThe project follows standard Go conventions:
- Use
gofmtfor formatting - Run
go vetto catch common issues - Use
golangci-lintfor comprehensive linting
- Uses tinySQL for embedded database
- Data persisted in
.gobformat - Three main stores:
- Chunks: Vector embeddings and text content
- Chats: Conversation history
- Sources: Document metadata
- Cosine similarity for semantic search
- Configurable chunk size and retrieval count (k)
- Efficient in-memory vector operations
- Native tinySQL
VEC_SEARCHwith optional bounded result caching and privacy-preserving query analytics - Portable database snapshots for backend-independent administrative backups
- Optional hybrid retrieval (
HYBRID_SEARCH: vector + BM25 via reciprocal rank fusion) and configurable ANN index (flat/ivf/hnsw) with startup warm-up (VEC_WARM) - Metadata-aware R³ ranking with trust, quality, freshness, feedback, and sensitivity penalties
- Structure-aware chunking: text is split into atomic blocks first — a run of numbered/bulleted list items, a table's rows, or a fenced code block — so a character-budget cut lands between blocks, not in the middle of a step-by-step procedure or a table row. Plain prose chunking behavior is unchanged. An oversized block/line that still exceeds the budget is now split by its own rows/lines (falling back to word-boundary slicing only for a genuinely unbreakable line), instead of passing through unbounded.
- Office document structure preservation: DOCX/PPTX/XLSX/ODT/ODP/ODS text extraction now emits a line break at paragraph, table row, and table cell boundaries (instead of collapsing an entire document/sheet into one line), so the structure-aware chunker above can recognize spreadsheet rows and table structure in these formats too.
- Source-type quality tiers (
r3.go) include a genericstructured_itemtier for small, well-formed structured records — a glossary entry, a reference-table row, a course module or Q/A pair — imported one record at a time (CSV/JSON row import uses this tier automatically); the dedicatedwikitier is now actually applied by the Wikipedia import endpoint instead of silently falling back to the generic default. - Operator-configurable terminology (
terminologysetting, managed viaGET/POST /api/settings/terminology): a table of term ↔ expansion pairs (e.g. an abbreviation and its full form, or a domain synonym pair). When a query matches one side, query expansion adds the other side(s) as additional retrieval/search variants — bridging vocabulary the embedding model alone might not know is equivalent. Empty by default. - Revision-aware re-ingestion:
/api/add-wiki,/api/add-url, and/api/add-textnow accept an optionalmetadataobject (same shape as the folder-import path), includingupdate_mode(skip(default) /upsert/replace), so a previously imported page/URL/text can be refreshed by content hash instead of being silently skipped forever.
tinyRAG includes a governed retrieval layer (R³: Ranked, Responsible, Retrieval):
- Retrieval units with source/ACL/sensitivity/provenance metadata
- Weighted deterministic ranking (
R3Score) beyond pure semantic similarity - ACL and role filtering before context assembly
- Citation-first context and answer constraints
- Policy-driven tool persistence (
transient_only,persistable_after_policy,never_persist) - Import job and audit-event telemetry tables for traceability
See:
docs/r3-architecture.mddocs/import-adapters.mddocs/request-lifecycle.mddocs/xml-tool-protocol.mddocs/agentic-tool-use.md
- Protocol-aware inference client (
auto, OpenAI-compatible/v1, or native Ollama/api) - OpenAI-compatible inference servers, including GopherLLM, RustyLLM, llama.cpp, LM Studio, vLLM and gateways
- Native Ollama model discovery, chat streaming, batched embeddings and legacy embedding fallback
- Streaming responses
- Support for custom system prompts (personas)
- Context injection from retrieved chunks
For machine-to-machine jobs you can use POST /api/process.
This endpoint is intended for workflows where a Python or PHP tool:
- reads rows from MSSQL
- sends each row or a grouped payload as JSON
- passes
system_prompt,pre_prompt, andpost_prompt - requests a strict JSON response with a provided schema
- validates the JSON on the server
- writes the result to JSONL
Example request:
{
"request_id": "row-4711",
"mode": "direct",
"system_prompt": "Du bist ein Extraktionssystem.",
"pre_prompt": "Analysiere die Eingabedaten.",
"input": {
"id": 4711,
"company": "Example GmbH",
"text": "..."
},
"post_prompt": "Gib nur JSON zurueck.",
"response_schema": {
"type": "object",
"required": ["status", "summary"],
"additionalProperties": false,
"properties": {
"status": {"type": "string"},
"summary": {"type": "string"},
"score": {"type": "number"}
}
},
"options": {
"validate_json": true,
"repair_json": true,
"max_retries": 2
}
}Example response:
{
"request_id": "row-4711",
"ok": true,
"mode": "direct",
"valid_json": true,
"attempts": 1,
"duration_ms": 812,
"raw": "{\"status\":\"ok\",\"summary\":\"...\",\"score\":0.94}",
"result": {
"status": "ok",
"summary": "...",
"score": 0.94
}
}If you want retrieval from the local knowledge base before processing, set mode to "rag" or rag.enabled to true.
Included example:
internal/app/examples/gallery.htmlis the embedded theme and scenario gallery served at/gallery.
By default, code execution features are disabled for security:
allow_code_exec: Allows running user-provided codeallow_nanogo: Enables nanoGo interpreter
- No built-in authentication
- Recommended to run behind a reverse proxy with auth
- Consider network isolation for production use
- github.com/SimonWaldherr/tinySQL - Embedded SQL database
- simonwaldherr.de/go/nanogo - Go interpreter
- simonwaldherr.de/go/smallr - Small templating engine
See the repository for license information.
Contributions are welcome! Please feel free to submit issues and pull requests.
Simon Waldherr - GitHub
- tinySQL - Lightweight embedded SQL database for Go