Skip to content

Tool selection precision: hallucinated tool calls on knowledge / RAG queries #209

Description

@pluginslab

Summary

Qwen 3 1.7B (prompt-based JSON mode) sometimes picks a clearly irrelevant tool when the user asks a knowledge question — including cases where the user has explicitly enabled the knowledge-base (book / `docSearch`) toggle. The model can route to `final_answer` instead, but currently doesn't reliably do so under weak tool signal.

This umbrella supersedes #207 and bundles two related precision fixes.

Reproductions

Case A — meta question, no docSearch (was #207):

do you have an ability that adds context to the page where you are calling this chat from?

Router → ReAct → LLM picks `wp-agentic-admin/read-file` and fails on missing path.

Case B — knowledge question with docSearch enabled (book button on):

Whats the hook that allows me to inject content in the footer?

Router → ReAct (book button only prepends RAG snippets to the message — it doesn't change routing). LLM picks `wp-agentic-admin/role-capabilities-check` despite zero keyword/description overlap.

Root cause

  • `src/extensions/services/message-router.js:103` has no awareness of `docSearch`. All paths that don't match a workflow end at `{ type: 'react' }` (steps 3 and 4 fall through to ReAct).
  • The book button (`ChatInput.jsx:307`) sets `docSearchEnabled`, which `react-agent.js:178` uses only to inject RAG snippets into the user message — tool selection still runs.
  • Under weak signal, Qwen 3 1.7B in prompt-based JSON mode grabs the closest-sounding tool instead of emitting `final_answer`.
  • Read-file's description ("Always use this tool for file reading requests") amplifies the bias toward over-selection (was Meta questions about capabilities get routed to tools (e.g. read-file) #207's specific case).

Proposed fixes

Fix 1 — Honor explicit RAG intent: skip ReAct when `docSearch` is on

When the book button is enabled and RAG returns hits, route directly to the conversational path with the augmented prompt. The user has stated their intent; don't second-guess by fishing for tools.

  • Plumb `docSearch` into `message-router.js#route()` (or short-circuit before it in `chat-orchestrator.js`).
  • Open question: what should happen if `docSearch` is on but the vector store returns zero hits? Probably still skip ReAct — user explicitly asked for a knowledge answer.

Fix 2 — Post-hoc validation of tool selection

After the LLM emits a tool choice, validate that at least one of the tool's keywords or description nouns appears in the user message. If not, treat as hallucination and force `final_answer`.

  • Add validation in `react-agent.js` around the JSON-parse of the LLM's response.
  • Cheap, deterministic, doesn't rely on the LLM behaving well.

Fix 3 (from #207, narrower) — Soften read-file priming

  • Drop "Always use this tool" from read-file's description.
  • Trim broad verbs from its keywords (`show`, `view`, `open`, `contents`, `source`).
  • Lowest-risk change of the three; ship even if 1 and 2 take longer.

Notes

Test plan

  • Case A reproduces conversational answer (no tool call)
  • Case B reproduces conversational answer using RAG context
  • Existing tool calls ("list plugins", "flush cache", etc.) still route correctly
  • Add unit test for post-hoc validation

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions