CanonicAI
The canonical-data factory

Corpus in. Canonical data out.

CanonicAI turns unstructured knowledge — books, research papers, domain documents — into canonical, queryable datasets. At scale, with provenance.

Commission a canon →

“Anyone can write a prompt. The defensible thing is running thousands of multi-step extractions reliably, idempotently, and with lineage — at a cost you control. That’s why factory is the honest metaphor: a production line with QA and inventory control, not a clever prompt.”

— THE OPERATING THESIS

The production lines

Book Factory

Whole books deconstructed at chapter-respecting fidelity into tagged summaries, argument models, and factor structures — not naive chunks.

Article Factory

Peer-reviewed research distilled into instruments, constructs, citations, and effect data — feeding a living evidence engine.

Compendium Factory

Reference catalogs of validated measurement scales extracted item-by-item — validated against known ground truth at 95% recall.

Schema Authority

One canonical measurement vocabulary — constructs, items, instruments, effect sizes — defined once, conformed to by every consumer.

The line, running

8,542
Registered assets
700+
Instruments extracted
4,900+
Books modeled
SHA-256
Provenance, every source

Canon catalog

Book Elements

Summaries, narrative structures, constructs, relationships, and source-grounded models across more than 4,900 works.

Built

Research & Evidence

Research assets resolved into studies, citations, effects, constructs, and reusable evidence records.

Built

Instrument Canon

Measurement instruments, scales, items, response options, scoring notes, and construct associations.

Built

Job Elements

Occupations, tasks, knowledge, skills, abilities, activities, styles, and interests grounded in O*NET.

Built

Metric Canon

Operational and scientific measures with definitions, entities, formulas, provenance, and deduplicated identities.

Built

Regulation & Statistics

Wage, labour-regulation, and labour-statistics records normalized for comparison and downstream computation.

Built

Competency Elements

Competency statements and classifications transformed into addressable, joinable elements.

Built

Taxonomy & Routing

Controlled vocabularies and routing canons that place heterogeneous source material into one navigable system.

Built
CATALOG CHECKED · SEP 2026
What the catalog means. These counts demonstrate production throughput; they are not source inventory offered for resale. A commissioned canon is produced from material the client owns, licenses, or provides, with the resulting rights and delivery terms defined for that engagement.

Frames — shared coordinate systems

A canon can be more than a list. A Frame gives every record an address in a shared coordinate system, so applications and agents can join, compare, translate, and compose data without bespoke integrations.

JobFrame — live SkillFrame — live GeoFrame — live ProblemFrame — planned PersonFrame — planned IndustryFrame — planned ConditionFrame — planned StoryFrame — planned

What a record looks like

The product is not a chat transcript. It is durable, joinable records with stable IDs, aliases, provenance, and typed relationships — the kind of object an application can query a year later.

{
  "canonical_id": "business_strategy_alignment",
  "name": "Business/Strategy Alignment of Analytics",
  "tier": "core",
  "four_s": "strategy",
  "aliases": ["Business Priority Alignment", "Problem Framing…"],
  "relationship": {
    "from": "business_strategy_alignment",
    "to": "analytics_capability",
    "type": "enables",
    "support": 0.14,
    "provenance": ["van_vulpen", "fundamentals_of_hr_analytics"]
  }
}
Excerpt from the people-analytics cluster model — 45 constructs · 45 relationships · book-level provenance on every edge
Scattered extraction

“Organizations struggling with unclear strategies can use the Diamond Model to define desired outcomes and develop corresponding competencies and capabilities.”

Joinable record
{
  "problem": "unclear_strategy",
  "solution": "diamond_model",
  "book": "rewarding_excellence",
  "chapter": "ch01",
  "source_kind": "problems_addressed"
}
Same meaning, now addressable: problem ↔ solution ↔ book ↔ chapter. The wording remains preserved in provenance.

Taxonomy as address space

Heterogeneous sources only become a catalog when every item receives a stable address. JobFrame alone carries 14,948 occupational cells; the guide taxonomy classifies thousands of works into navigable tracks.

Function
wf_nursing · Nursing
Family
jf_inpatient · bedside care in hospitals
jf_outpatient · clinic and ambulatory care
jf_long_term_care · extended residence and memory care
jf_in_home_care · home nursing and hospice
Guide track
people-analytics · lead-and-manage · analyze-compensation · run-a-survey · earn-a-phd · get-published · 48 tracks over 2,541 classified works
Four-S
science · statistics · systems · strategy — the shared analytic spine every construct is placed on
People analytics construct relationships Business strategy alignment, data quality, team skillset, tools, and leadership enable analytics capability, which enables evidence-based decisions and produces HR credibility. ENABLES ENABLES PRODUCES Strategy alignment support .14 · 3 books Data quality support .29 · 7 books Team skillset support .14 · 3 books Tools & technology support .11 · 3 books Leadership sponsorship support .11 · 3 books Analytics capability canonical construct Evidence-based decisions HR credibility & influence
A real subgraph from the people-analytics model. Purple edges are core; every support score is backed by named books.

How a problem gets resolved

Problems already hide in the corpus as attributes on tools, brandscripts, deep extracts, and cases. Making them first-class is harder than it sounds: statements conflict, synonyms proliferate, and every resolution has to carry provenance or it cannot be trusted.

01

Scatter

Problems live as strings on tools, frameworks, brandscripts, deep extracts, and cases — never as records. Useful locally; unjoinable globally.

Failure mode: treat every string as a unique problem → catalog explosion
02

Harvest

Deterministic lift with full provenance. One cluster prove-one already produced 514 raw statements across 9 books — 435 from chapter problems_addressed alone — at $0.

Hard part: keep verbatim text; count duplicates honestly; invent nothing
03

Resolve

Cluster and dedupe into canonical Problem Cards. Measure rater agreement; withhold low-agreement merges. A synonym map without evidence is fiction.

Hard part: “retention” ≠ “regretted attrition of high performers” — hierarchy, not collapse
04

Join

Wire each card to solutions, models, personas, jobs, and stories. The differentiator is diagnosis first, then better-fitting solutions grounded in evidence.

Hard part: brandscripts are entity×problem joins, not one-problem-per-entity attributes
05

Surface

Browse-by-problem, wizards that diagnose before they recommend, guides whose Orient movement opens on a real Problem Card — not a marketing slogan.

Hard part: every consumer must speak the same IDs, or the net tears

Wide nets from one canon

Once the factory produces a guide, model, or occupational cell, many surfaces can publish from it. Same records. Different fronts. That is the point of a canon.

People analytics — destination guide

Full depth on the PeopleAnalyst surface: constructs, playbook, measurement joins.

peopleanalyst.com ↗

Same subject — Bicycle front

The travel-guide rendering of the same corpus model: route, stops, honest disagreement.

bicycle.guide ↗

Occupational spine — JobFrame

14,948 cells as a shared address bus for careers, compensation, and capability joins.

jobframe.global ↗

Domain home — compensation

CompensationProfessional consumes the same producer surface for a specialist audience.

compensationprofessional.com ↗

Reference library

The shared book corpus as a navigable library — identity and status, not a second catalog.

peopleanalyst.com/library ↗

Analytics that need a canon

A single PDF or a vector store can answer a question. A canon answers questions that cross books, instruments, jobs, and evidence — questions that are otherwise impossible without months of hand joins.

Construct coverage across a field

Which capabilities appear as core across dozens of people-analytics books, which are contested, and which edges are weakly supported? That is a graph query over stable construct IDs — not a keyword search.

Requires: cluster models · durable IDs · typed relationships · book provenance

Problem → solution inversion

Start from “unclear strategy / weak incentive design” and retrieve every framework that claims to address it, with chapter provenance. One cluster already yields 435 such joins before resolution.

Requires: harvest of problems_addressed · solution names · chapter paths

Instrument × construct × effect

Which validated scales measure which constructs, and what effect sizes appear where? Evidence products need item-level instrument canons joined to construct space — not abstracts glued to PDFs.

Requires: instrument canon · construct registry · study/effect extracts

Occupation × capability joins

Map a capability model onto JobFrame cells to ask which roles need which constructs, or which guides should bind to which jobs. Impossible without a shared occupational address bus.

Requires: JobFrame / SOC spine · guide↔job bindings · capability models

Sellability from measured depth

Price and gate guides from counted paid bodies, not vibes. Empty rooms and free projections become operational signals the factory can act on.

Requires: depth sidecars · entitlement contracts · supply accounting
Instrument canon · example record

Trust Scale — faith and confidence in peers and management

12 items · 7-point Likert · 6 covered constructs · source-addressed

“If I got into difficulties at work I know my workmates would try and help me out.”

“Management at work seems to do an efficient job.”

“Our management would be quite prepared to gain advantage by deceiving the workers.” (reverse-scored)

1strongly disagree
2disagree quite a lot
3disagree a little
4not sure
5agree a little
6agree quite a lot
7strongly agree
Canonical output preserves item text, response anchors, reverse scoring, constructs, source asset ID, and extraction provenance.

Commission a canon

Prove one unit before you fund the line.

Bring a standards library, scientific literature set, technical-documentation corpus, historical archive, or proprietary knowledge base. We scope one representative unit, run it end to end, show the canonical output and provenance trail, and measure its real production cost before proposing scale.

Tell us about your material.

Your pilot brief.

This is the starting point for the scoping conversation — it's included in the email you send below.

Pilot brief

The 30-min call settles
  • Material fit and acquisition path
  • Per-unit production cost and timeline
  • Delivery format for the canonical output

Who should we confirm the call with?

✓

Your pilot brief is ready.

Send the brief below, then book the 30-minute scoping conversation.

Glass box, not black box

Every dataset CanonicAI ships traces to its source — file hashes, extraction lineage, model and prompt provenance, idempotent re-derivation. If a number is in the output, you can walk it back to the page it came from.

The engine is the producer and source-of-truth; everything downstream is a consumer. It powers the PeopleAnalyst family:

PeopleAnalyst — the destination Principia — the evidence engine People Analytics Toolbox