Skip to content

Improve embedding fidelity - #1449

Merged
malatewang merged 6 commits into
MemMachine:mainfrom
edwinyyyu:embedding_chunking_fidelity
Sep 18, 2026
Merged

malatewang merged 6 commits into
MemMachine:mainfrom
edwinyyyu:embedding_chunking_fidelity

Conversation

@edwinyyyu

@edwinyyyu edwinyyyu commented Jun 11, 2026 •

Copy link
Copy Markdown
Contributor

Purpose of the change

Chunk-merged embeddings (texts longer than the model input limit are split, embedded per chunk, and averaged) reproduce the whole-text embedding more faithfully with greedy max-size chunks and a length-weighted average than with the current balanced chunks and unweighted mean. Measured on three encoders; numbers below.

The special-token sanitizer in the OpenAI embedder is removed. It was a workaround for a server-side bug at api.openai.com (text-embedding-3-small returned 500 on inputs containing tokens such as <|endoftext|>), and that bug is gone: 7 tokens x 3 models x 3 input forms all return 200 with valid embeddings (probe script and results attached as a PR comment below). The workaround was not free: it deleted the literal token from the text before embedding, so a stored transcript that quotes <|endoftext|> was embedded as if it did not. It is also OpenAI-specific by design. OpenAIEmbedder is also used with OpenAI-compatible endpoints (Ollama, DashScope), which the probe does not cover; an endpoint that fails on a plain string containing such a token is broken at the endpoint, and the remedy is to not use it, not for MemMachine to pre-edit input for every provider that might mis-tokenize.

max_input_length above the OpenAI embedder's per-request cap (75,000 code points) is now clamped to the cap. A limit above the cap can never be honored (no chunk larger than the cap fits in any request), and before this change it reached cluster_texts with an oversized chunk and raised ValueError (HTTP 500), the #1298 failure for that configuration. Greedy chunking widened which input lengths hit it: at max_input_length=80_000, balanced chunking crashed on 76,000-80,000 and 160,000-char inputs; greedy crashed on everything above 75,000. No shipped sample config sets a limit above the cap.

Description

  • Use greedy max-size chunks (chunk_text) instead of balanced chunks (chunk_text_balanced) in the OpenAI, SentenceTransformer, and Amazon Bedrock embedders.
  • Merge chunk embeddings with a length-weighted np.average instead of unweighted np.mean.
  • Remove the special-token sanitizer from the OpenAI embedder; keep the empty-input coercion (or "."), which is independently load-bearing: the API rejects empty strings and chunk_text("") yields no chunks to average.
  • Clamp max_input_length to max_total_input_length_per_request in the OpenAI embedder, with a regression test that fails before the clamp (ValueError from cluster_texts) and passes after.

Measurements

Method: 10 hand-written texts in distinct registers (fiction, technical docs, chat transcript, news, legal, academic, recipe, business memo, product review, encyclopedia), each short enough to embed whole — the whole-text embedding is the observable ground truth. Each text is chunked at limits 800 / 1600 / 3200 chars under each policy, chunks embedded, merged, and compared to the whole-text vector by cosine similarity. Single-chunk cells excluded; 28 cells per configuration per model. Probe script and raw per-cell results are attached as PR comments below (deliberately not committed to the branch).

This PR's configuration (greedy + length-weighted) vs current behavior (balanced + unweighted), paired over the same 28 cells:

Encoder Current mean cos PR mean cos Paired Δ Cell wins t (df=27)
embeddinggemma-300m (local) 0.9578 0.9644 +0.0066 24/28 4.8
text-embedding-3-small 0.9338 0.9394 +0.0055 18/28 2.3
text-embedding-3-large 0.9237 0.9340 +0.0103 20/28 3.3

Full split × weighting grid (mean cosine to whole-text embedding, higher = more faithful):

Configuration gemma-300m 3-small 3-large
greedy + length-weighted (this PR) 0.9644 0.9394 0.9340
balanced + weighted 0.9578 0.9338 0.9237
balanced + unweighted (current) 0.9578 0.9338 0.9237
greedy + unweighted 0.9356 0.9217 0.9161

Two structural checks that this is signal, not noise:

  • Greedy requires the weighting: greedy+unweighted ranks last on all three encoders. The two changes in this PR are a package.
  • Balanced is indifferent to weighting (its two rows agree to ~5 decimals on every encoder), as expected for equal-size chunks — a built-in sanity check on the harness.

Mechanism: to the extent an encoder mean-pools token states, the whole-text vector is approximately a length-weighted average over the text, which a merge approximates best when weights match chunk lengths and most characters sit in maximal-context chunks. The result holds on the OpenAI models even though their pooling is undocumented.

Caveat: per-cell variance is wider on the OpenAI models than on gemma (worst cell −0.019, best +0.048), so this is a consistent aggregate win rather than a uniform per-text one.

Type of change

  • Bug fix (non-breaking change which fixes an issue)

How Has This Been Tested?

  • Unit Test
  • Integration Test
  • End-to-end Test
  • Test Script (please provide)
  • Manual verification (list step-by-step instructions)

Test Results: Fidelity tables above; probe scripts and raw results attached as PR comments below; embedder unit tests pass (server_tests/memmachine_server/common/embedder/). The fidelity change itself is not pinned by a unit test: its contract (closeness to the whole-text embedding) is only observable against a real encoder, and the measurement above is the evidence for it.

Checklist

  • I have signed the commit(s) within this pull request
  • My code follows the style guidelines of this project (See STYLE_GUIDE.md)
  • I have performed a self-review of my own code
  • I have commented my code
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added unit tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes
  • Any dependent changes have been merged and published in downstream modules
  • I have checked my code and corrected any misspellings

Maintainer Checklist

  • Confirmed all checks passed
  • Contributor has signed the commit(s)
  • Reviewed the code
  • Run, Tested, and Verified the change(s) work as expected

@edwinyyyu

Copy link
Copy Markdown
Contributor Author

Probe attachment 1/2 — chunking fidelity (kept out of the branch deliberately; attached for reviewers)

probe_chunking_fidelity.py
"""Probe: does the greedy+length-weighted chunk-merge win hold on OpenAI embedders?

The EmbeddingGemma-300m experiment (memlite, docs/design/23) found that for
texts split-embedded-merged, greedy max-size chunks with a length-weighted
average reproduce the whole-text embedding more faithfully than balanced
chunks with an unweighted average (the previous memmachine behavior). This
probe reruns that comparison against OpenAI's embedding models, using the
same ten register texts (parsed out of the memlite experiment harness so the
corpora stay identical).

Configurations (split x weighting):
  exact+greedy   — chunk_text: max-size chunks, exact cuts (new behavior)
  exact+balanced — chunk_text_balanced: even ~len/n chunks (old behavior)
  weighted       — np.average weighted by chunk length (new behavior)
  unweighted     — plain np.mean (old behavior)

Boundary-snapped variants are omitted: snapping has never existed in the
Python embedders, so it is not a decision-relevant axis here.

Run: uv run python probe_chunking_fidelity.py
"""

import json
import math
import re
import sys
from collections import defaultdict
from pathlib import Path

import numpy as np
import openai

ENV_PATH = Path(
    "/Users/eyu/edwinyyyu/mmcc/segment_store/evaluation/event_memory/locomo/.env"
)
RUST_TEXTS_PATH = Path(
    "/Users/eyu/edwinyyyu/memlite/perf/crates/core/src/service/tests/embedding_fidelity.rs"
)
RESULTS_PATH = Path(__file__).with_suffix(".results.json")

MODELS = [
    "text-embedding-3-small",
    "text-embedding-3-large",
]

CHUNK_LIMITS = [800, 1600, 3200]


def load_api_key() -> str:
    for line in ENV_PATH.read_text().splitlines():
        line = line.strip()
        if line.startswith("OPENAI_API_KEY"):
            return line.split("=", 1)[1].strip().strip('"').strip("'")
    raise RuntimeError(f"OPENAI_API_KEY not found in {ENV_PATH}")


def load_register_texts() -> dict[str, str]:
    """Parse the `const NAME: &str = "\\` ... `";` texts from the Rust harness."""
    texts: dict[str, str] = {}
    name = None
    lines: list[str] = []
    for line in RUST_TEXTS_PATH.read_text().splitlines():
        if name is None:
            match = re.match(r'const (\w+): &str = "\\$', line.strip())
            if match:
                name = match.group(1).lower()
                lines = []
            continue
        if line.endswith('";'):
            lines.append(line[: -len('";')])
            texts[name] = "\n".join(lines)
            name = None
        else:
            lines.append(line)
    if len(texts) != 10:
        raise RuntimeError(f"expected 10 register texts, parsed {len(texts)}")
    return texts


def chunk_greedy(text: str, max_length: int) -> list[str]:
    return [text[i : i + max_length] for i in range(0, len(text), max_length)]


def chunk_balanced(text: str, max_length: int) -> list[str]:
    num_chunks = math.ceil(len(text) / max_length)
    chunk_size = math.ceil(len(text) / num_chunks)
    return [text[i : i + chunk_size] for i in range(0, len(text), chunk_size)]


def embed_batch(client: openai.OpenAI, model: str, texts: list[str]) -> list[np.ndarray]:
    response = client.embeddings.create(model=model, input=texts)
    ordered = sorted(response.data, key=lambda item: item.index)
    if len(ordered) != len(texts):
        raise RuntimeError(f"{model}: got {len(ordered)} vectors for {len(texts)} inputs")
    return [np.asarray(item.embedding, dtype=float) for item in ordered]


def cosine(a: np.ndarray, b: np.ndarray) -> float:
    return float(np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b)))


def main() -> int:
    client = openai.OpenAI(api_key=load_api_key())
    texts = load_register_texts()

    results = []
    for model in MODELS:
        totals: dict[str, list[float]] = defaultdict(list)
        wins: dict[str, int] = defaultdict(int)
        for register, text in texts.items():
            whole = embed_batch(client, model, [text])[0]
            for limit in CHUNK_LIMITS:
                splits = {
                    "exact+greedy": chunk_greedy(text, limit),
                    "exact+balanced": chunk_balanced(text, limit),
                }
                if any(len(chunks) < 2 for chunks in splits.values()):
                    continue
                cell: dict[str, float] = {}
                for split_name, chunks in splits.items():
                    vectors = embed_batch(client, model, chunks)
                    weights = [len(chunk) for chunk in chunks]
                    merged = {
                        "weighted": np.average(vectors, axis=0, weights=weights),
                        "unweighted": np.mean(vectors, axis=0),
                    }
                    for merge_name, vector in merged.items():
                        label = f"{split_name}/{merge_name}"
                        cos = cosine(vector, whole)
                        cell[label] = cos
                        totals[label].append(cos)
                        results.append(
                            {
                                "model": model,
                                "register": register,
                                "chunk_limit": limit,
                                "config": label,
                                "chunks": len(splits[split_name]),
                                "cosine": cos,
                            }
                        )
                best = max(cell, key=cell.get)
                wins[best] += 1
                print(
                    f"{model} {register:>13} limit={limit:>5} "
                    + " ".join(f"{k}={v:.5f}" for k, v in sorted(cell.items()))
                )

        print(f"\n=== {model}: mean cosine to whole-text embedding ===")
        ranked = sorted(
            ((np.mean(v), k, len(v)) for k, v in totals.items()), reverse=True
        )
        for mean, label, count in ranked:
            print(
                f"  {label:<26} mean_cos={mean:.5f} "
                f"cell_wins={wins.get(label, 0)} n={count}"
            )

        # Paired comparison: new behavior vs old behavior.
        new = [r["cosine"] for r in results if r["model"] == model and r["config"] == "exact+greedy/weighted"]
        old = [r["cosine"] for r in results if r["model"] == model and r["config"] == "exact+balanced/unweighted"]
        deltas = [n - o for n, o in zip(new, old, strict=True)]
        mean_delta = float(np.mean(deltas))
        sd = float(np.std(deltas, ddof=1))
        t_stat = mean_delta / (sd / math.sqrt(len(deltas))) if sd > 0 else float("inf")
        print(
            f"  paired greedy/weighted vs balanced/unweighted: "
            f"delta={mean_delta:+.5f} wins={sum(d > 0 for d in deltas)}/{len(deltas)} t={t_stat:.2f}"
        )

    RESULTS_PATH.write_text(json.dumps(results, indent=2))
    print(f"\nwrote {RESULTS_PATH}")
    return 0


if __name__ == "__main__":
    sys.exit(main())
probe_chunking_fidelity.results.json (per-cell raw data, 224 rows)
[
  {
    "model": "text-embedding-3-small",
    "register": "fiction",
    "chunk_limit": 800,
    "config": "exact+greedy/weighted",
    "chunks": 7,
    "cosine": 0.8827412434172709
  },
  {
    "model": "text-embedding-3-small",
    "register": "fiction",
    "chunk_limit": 800,
    "config": "exact+greedy/unweighted",
    "chunks": 7,
    "cosine": 0.8609035488024998
  },
  {
    "model": "text-embedding-3-small",
    "register": "fiction",
    "chunk_limit": 800,
    "config": "exact+balanced/weighted",
    "chunks": 7,
    "cosine": 0.8803453623398728
  },
  {
    "model": "text-embedding-3-small",
    "register": "fiction",
    "chunk_limit": 800,
    "config": "exact+balanced/unweighted",
    "chunks": 7,
    "cosine": 0.8803837930862868
  },
  {
    "model": "text-embedding-3-small",
    "register": "fiction",
    "chunk_limit": 1600,
    "config": "exact+greedy/weighted",
    "chunks": 4,
    "cosine": 0.9398753780348335
  },
  {
    "model": "text-embedding-3-small",
    "register": "fiction",
    "chunk_limit": 1600,
    "config": "exact+greedy/unweighted",
    "chunks": 4,
    "cosine": 0.8917911316257472
  },
  {
    "model": "text-embedding-3-small",
    "register": "fiction",
    "chunk_limit": 1600,
    "config": "exact+balanced/weighted",
    "chunks": 4,
    "cosine": 0.9374298961547437
  },
  {
    "model": "text-embedding-3-small",
    "register": "fiction",
    "chunk_limit": 1600,
    "config": "exact+balanced/unweighted",
    "chunks": 4,
    "cosine": 0.9374457347638924
  },
  {
    "model": "text-embedding-3-small",
    "register": "fiction",
    "chunk_limit": 3200,
    "config": "exact+greedy/weighted",
    "chunks": 2,
    "cosine": 0.9733366810256934
  },
  {
    "model": "text-embedding-3-small",
    "register": "fiction",
    "chunk_limit": 3200,
    "config": "exact+greedy/unweighted",
    "chunks": 2,
    "cosine": 0.9641443896322912
  },
  {
    "model": "text-embedding-3-small",
    "register": "fiction",
    "chunk_limit": 3200,
    "config": "exact+balanced/weighted",
    "chunks": 2,
    "cosine": 0.973897249035695
  },
  {
    "model": "text-embedding-3-small",
    "register": "fiction",
    "chunk_limit": 3200,
    "config": "exact+balanced/unweighted",
    "chunks": 2,
    "cosine": 0.973897249035695
  },
  {
    "model": "text-embedding-3-small",
    "register": "technical",
    "chunk_limit": 800,
    "config": "exact+greedy/weighted",
    "chunks": 7,
    "cosine": 0.8702421041249913
  },
  {
    "model": "text-embedding-3-small",
    "register": "technical",
    "chunk_limit": 800,
    "config": "exact+greedy/unweighted",
    "chunks": 7,
    "cosine": 0.8606649205422204
  },
  {
    "model": "text-embedding-3-small",
    "register": "technical",
    "chunk_limit": 800,
    "config": "exact+balanced/weighted",
    "chunks": 7,
    "cosine": 0.8735199472689553
  },
  {
    "model": "text-embedding-3-small",
    "register": "technical",
    "chunk_limit": 800,
    "config": "exact+balanced/unweighted",
    "chunks": 7,
    "cosine": 0.8735598793420628
  },
  {
    "model": "text-embedding-3-small",
    "register": "technical",
    "chunk_limit": 1600,
    "config": "exact+greedy/weighted",
    "chunks": 4,
    "cosine": 0.937736362909748
  },
  {
    "model": "text-embedding-3-small",
    "register": "technical",
    "chunk_limit": 1600,
    "config": "exact+greedy/unweighted",
    "chunks": 4,
    "cosine": 0.9063515882853833
  },
  {
    "model": "text-embedding-3-small",
    "register": "technical",
    "chunk_limit": 1600,
    "config": "exact+balanced/weighted",
    "chunks": 4,
    "cosine": 0.9003253906128421
  },
  {
    "model": "text-embedding-3-small",
    "register": "technical",
    "chunk_limit": 1600,
    "config": "exact+balanced/unweighted",
    "chunks": 4,
    "cosine": 0.900332353994952
  },
  {
    "model": "text-embedding-3-small",
    "register": "technical",
    "chunk_limit": 3200,
    "config": "exact+greedy/weighted",
    "chunks": 2,
    "cosine": 0.9737398691098803
  },
  {
    "model": "text-embedding-3-small",
    "register": "technical",
    "chunk_limit": 3200,
    "config": "exact+greedy/unweighted",
    "chunks": 2,
    "cosine": 0.9693985987839949
  },
  {
    "model": "text-embedding-3-small",
    "register": "technical",
    "chunk_limit": 3200,
    "config": "exact+balanced/weighted",
    "chunks": 2,
    "cosine": 0.9610261061239891
  },
  {
    "model": "text-embedding-3-small",
    "register": "technical",
    "chunk_limit": 3200,
    "config": "exact+balanced/unweighted",
    "chunks": 2,
    "cosine": 0.9610290742498959
  },
  {
    "model": "text-embedding-3-small",
    "register": "chat",
    "chunk_limit": 800,
    "config": "exact+greedy/weighted",
    "chunks": 5,
    "cosine": 0.8902357093789164
  },
  {
    "model": "text-embedding-3-small",
    "register": "chat",
    "chunk_limit": 800,
    "config": "exact+greedy/unweighted",
    "chunks": 5,
    "cosine": 0.8994408283753803
  },
  {
    "model": "text-embedding-3-small",
    "register": "chat",
    "chunk_limit": 800,
    "config": "exact+balanced/weighted",
    "chunks": 5,
    "cosine": 0.8853332978510506
  },
  {
    "model": "text-embedding-3-small",
    "register": "chat",
    "chunk_limit": 800,
    "config": "exact+balanced/unweighted",
    "chunks": 5,
    "cosine": 0.8853872203046886
  },
  {
    "model": "text-embedding-3-small",
    "register": "chat",
    "chunk_limit": 1600,
    "config": "exact+greedy/weighted",
    "chunks": 3,
    "cosine": 0.9218884956021454
  },
  {
    "model": "text-embedding-3-small",
    "register": "chat",
    "chunk_limit": 1600,
    "config": "exact+greedy/unweighted",
    "chunks": 3,
    "cosine": 0.9358280578079553
  },
  {
    "model": "text-embedding-3-small",
    "register": "chat",
    "chunk_limit": 1600,
    "config": "exact+balanced/weighted",
    "chunks": 3,
    "cosine": 0.9285699309130073
  },
  {
    "model": "text-embedding-3-small",
    "register": "chat",
    "chunk_limit": 1600,
    "config": "exact+balanced/unweighted",
    "chunks": 3,
    "cosine": 0.9285892198662347
  },
  {
    "model": "text-embedding-3-small",
    "register": "chat",
    "chunk_limit": 3200,
    "config": "exact+greedy/weighted",
    "chunks": 2,
    "cosine": 0.9473384997376371
  },
  {
    "model": "text-embedding-3-small",
    "register": "chat",
    "chunk_limit": 3200,
    "config": "exact+greedy/unweighted",
    "chunks": 2,
    "cosine": 0.9455348808823826
  },
  {
    "model": "text-embedding-3-small",
    "register": "chat",
    "chunk_limit": 3200,
    "config": "exact+balanced/weighted",
    "chunks": 2,
    "cosine": 0.9489477662246923
  },
  {
    "model": "text-embedding-3-small",
    "register": "chat",
    "chunk_limit": 3200,
    "config": "exact+balanced/unweighted",
    "chunks": 2,
    "cosine": 0.9489477662246923
  },
  {
    "model": "text-embedding-3-small",
    "register": "news",
    "chunk_limit": 800,
    "config": "exact+greedy/weighted",
    "chunks": 7,
    "cosine": 0.922228616945852
  },
  {
    "model": "text-embedding-3-small",
    "register": "news",
    "chunk_limit": 800,
    "config": "exact+greedy/unweighted",
    "chunks": 7,
    "cosine": 0.9139745112678112
  },
  {
    "model": "text-embedding-3-small",
    "register": "news",
    "chunk_limit": 800,
    "config": "exact+balanced/weighted",
    "chunks": 7,
    "cosine": 0.9165243075196282
  },
  {
    "model": "text-embedding-3-small",
    "register": "news",
    "chunk_limit": 800,
    "config": "exact+balanced/unweighted",
    "chunks": 7,
    "cosine": 0.9165294508490657
  },
  {
    "model": "text-embedding-3-small",
    "register": "news",
    "chunk_limit": 1600,
    "config": "exact+greedy/weighted",
    "chunks": 4,
    "cosine": 0.9591208021097309
  },
  {
    "model": "text-embedding-3-small",
    "register": "news",
    "chunk_limit": 1600,
    "config": "exact+greedy/unweighted",
    "chunks": 4,
    "cosine": 0.9255461898560902
  },
  {
    "model": "text-embedding-3-small",
    "register": "news",
    "chunk_limit": 1600,
    "config": "exact+balanced/weighted",
    "chunks": 4,
    "cosine": 0.9496664274803505
  },
  {
    "model": "text-embedding-3-small",
    "register": "news",
    "chunk_limit": 1600,
    "config": "exact+balanced/unweighted",
    "chunks": 4,
    "cosine": 0.949670264919023
  },
  {
    "model": "text-embedding-3-small",
    "register": "news",
    "chunk_limit": 3200,
    "config": "exact+greedy/weighted",
    "chunks": 2,
    "cosine": 0.9812818248682762
  },
  {
    "model": "text-embedding-3-small",
    "register": "news",
    "chunk_limit": 3200,
    "config": "exact+greedy/unweighted",
    "chunks": 2,
    "cosine": 0.9732354569763894
  },
  {
    "model": "text-embedding-3-small",
    "register": "news",
    "chunk_limit": 3200,
    "config": "exact+balanced/weighted",
    "chunks": 2,
    "cosine": 0.9812410050566128
  },
  {
    "model": "text-embedding-3-small",
    "register": "news",
    "chunk_limit": 3200,
    "config": "exact+balanced/unweighted",
    "chunks": 2,
    "cosine": 0.9812401928545071
  },
  {
    "model": "text-embedding-3-small",
    "register": "legal",
    "chunk_limit": 800,
    "config": "exact+greedy/weighted",
    "chunks": 8,
    "cosine": 0.9015481184945144
  },
  {
    "model": "text-embedding-3-small",
    "register": "legal",
    "chunk_limit": 800,
    "config": "exact+greedy/unweighted",
    "chunks": 8,
    "cosine": 0.8962215555296424
  },
  {
    "model": "text-embedding-3-small",
    "register": "legal",
    "chunk_limit": 800,
    "config": "exact+balanced/weighted",
    "chunks": 8,
    "cosine": 0.9000880865655092
  },
  {
    "model": "text-embedding-3-small",
    "register": "legal",
    "chunk_limit": 800,
    "config": "exact+balanced/unweighted",
    "chunks": 8,
    "cosine": 0.900129671317358
  },
  {
    "model": "text-embedding-3-small",
    "register": "legal",
    "chunk_limit": 1600,
    "config": "exact+greedy/weighted",
    "chunks": 4,
    "cosine": 0.9387421059932247
  },
  {
    "model": "text-embedding-3-small",
    "register": "legal",
    "chunk_limit": 1600,
    "config": "exact+greedy/unweighted",
    "chunks": 4,
    "cosine": 0.9443071756200808
  },
  {
    "model": "text-embedding-3-small",
    "register": "legal",
    "chunk_limit": 1600,
    "config": "exact+balanced/weighted",
    "chunks": 4,
    "cosine": 0.9452211464066856
  },
  {
    "model": "text-embedding-3-small",
    "register": "legal",
    "chunk_limit": 1600,
    "config": "exact+balanced/unweighted",
    "chunks": 4,
    "cosine": 0.9452392525787787
  },
  {
    "model": "text-embedding-3-small",
    "register": "legal",
    "chunk_limit": 3200,
    "config": "exact+greedy/weighted",
    "chunks": 2,
    "cosine": 0.9786867115883692
  },
  {
    "model": "text-embedding-3-small",
    "register": "legal",
    "chunk_limit": 3200,
    "config": "exact+greedy/unweighted",
    "chunks": 2,
    "cosine": 0.9777571984609963
  },
  {
    "model": "text-embedding-3-small",
    "register": "legal",
    "chunk_limit": 3200,
    "config": "exact+balanced/weighted",
    "chunks": 2,
    "cosine": 0.9811498936246533
  },
  {
    "model": "text-embedding-3-small",
    "register": "legal",
    "chunk_limit": 3200,
    "config": "exact+balanced/unweighted",
    "chunks": 2,
    "cosine": 0.9811498936246533
  },
  {
    "model": "text-embedding-3-small",
    "register": "academic",
    "chunk_limit": 800,
    "config": "exact+greedy/weighted",
    "chunks": 5,
    "cosine": 0.8968408833368517
  },
  {
    "model": "text-embedding-3-small",
    "register": "academic",
    "chunk_limit": 800,
    "config": "exact+greedy/unweighted",
    "chunks": 5,
    "cosine": 0.8968189570565392
  },
  {
    "model": "text-embedding-3-small",
    "register": "academic",
    "chunk_limit": 800,
    "config": "exact+balanced/weighted",
    "chunks": 5,
    "cosine": 0.9108972734778836
  },
  {
    "model": "text-embedding-3-small",
    "register": "academic",
    "chunk_limit": 800,
    "config": "exact+balanced/unweighted",
    "chunks": 5,
    "cosine": 0.910883748437438
  },
  {
    "model": "text-embedding-3-small",
    "register": "academic",
    "chunk_limit": 1600,
    "config": "exact+greedy/weighted",
    "chunks": 3,
    "cosine": 0.9482695850537134
  },
  {
    "model": "text-embedding-3-small",
    "register": "academic",
    "chunk_limit": 1600,
    "config": "exact+greedy/unweighted",
    "chunks": 3,
    "cosine": 0.9297350232612078
  },
  {
    "model": "text-embedding-3-small",
    "register": "academic",
    "chunk_limit": 1600,
    "config": "exact+balanced/weighted",
    "chunks": 3,
    "cosine": 0.9395543661119989
  },
  {
    "model": "text-embedding-3-small",
    "register": "academic",
    "chunk_limit": 1600,
    "config": "exact+balanced/unweighted",
    "chunks": 3,
    "cosine": 0.9395300903322017
  },
  {
    "model": "text-embedding-3-small",
    "register": "academic",
    "chunk_limit": 3200,
    "config": "exact+greedy/weighted",
    "chunks": 2,
    "cosine": 0.9905572082575398
  },
  {
    "model": "text-embedding-3-small",
    "register": "academic",
    "chunk_limit": 3200,
    "config": "exact+greedy/unweighted",
    "chunks": 2,
    "cosine": 0.922136224208719
  },
  {
    "model": "text-embedding-3-small",
    "register": "academic",
    "chunk_limit": 3200,
    "config": "exact+balanced/weighted",
    "chunks": 2,
    "cosine": 0.9604994585571137
  },
  {
    "model": "text-embedding-3-small",
    "register": "academic",
    "chunk_limit": 3200,
    "config": "exact+balanced/unweighted",
    "chunks": 2,
    "cosine": 0.9604994585571137
  },
  {
    "model": "text-embedding-3-small",
    "register": "recipe",
    "chunk_limit": 800,
    "config": "exact+greedy/weighted",
    "chunks": 4,
    "cosine": 0.8921175243412754
  },
  {
    "model": "text-embedding-3-small",
    "register": "recipe",
    "chunk_limit": 800,
    "config": "exact+greedy/unweighted",
    "chunks": 4,
    "cosine": 0.8929430499598484
  },
  {
    "model": "text-embedding-3-small",
    "register": "recipe",
    "chunk_limit": 800,
    "config": "exact+balanced/weighted",
    "chunks": 4,
    "cosine": 0.8931213059042252
  },
  {
    "model": "text-embedding-3-small",
    "register": "recipe",
    "chunk_limit": 800,
    "config": "exact+balanced/unweighted",
    "chunks": 4,
    "cosine": 0.8931263978443825
  },
  {
    "model": "text-embedding-3-small",
    "register": "recipe",
    "chunk_limit": 1600,
    "config": "exact+greedy/weighted",
    "chunks": 2,
    "cosine": 0.9596141400065044
  },
  {
    "model": "text-embedding-3-small",
    "register": "recipe",
    "chunk_limit": 1600,
    "config": "exact+greedy/unweighted",
    "chunks": 2,
    "cosine": 0.957802769197888
  },
  {
    "model": "text-embedding-3-small",
    "register": "recipe",
    "chunk_limit": 1600,
    "config": "exact+balanced/weighted",
    "chunks": 2,
    "cosine": 0.9577344674052668
  },
  {
    "model": "text-embedding-3-small",
    "register": "recipe",
    "chunk_limit": 1600,
    "config": "exact+balanced/unweighted",
    "chunks": 2,
    "cosine": 0.9577164552152183
  },
  {
    "model": "text-embedding-3-small",
    "register": "memo",
    "chunk_limit": 800,
    "config": "exact+greedy/weighted",
    "chunks": 5,
    "cosine": 0.8909836515876899
  },
  {
    "model": "text-embedding-3-small",
    "register": "memo",
    "chunk_limit": 800,
    "config": "exact+greedy/unweighted",
    "chunks": 5,
    "cosine": 0.8870611864818144
  },
  {
    "model": "text-embedding-3-small",
    "register": "memo",
    "chunk_limit": 800,
    "config": "exact+balanced/weighted",
    "chunks": 5,
    "cosine": 0.9063920307965727
  },
  {
    "model": "text-embedding-3-small",
    "register": "memo",
    "chunk_limit": 800,
    "config": "exact+balanced/unweighted",
    "chunks": 5,
    "cosine": 0.9063548099358474
  },
  {
    "model": "text-embedding-3-small",
    "register": "memo",
    "chunk_limit": 1600,
    "config": "exact+greedy/weighted",
    "chunks": 3,
    "cosine": 0.9529238326048389
  },
  {
    "model": "text-embedding-3-small",
    "register": "memo",
    "chunk_limit": 1600,
    "config": "exact+greedy/unweighted",
    "chunks": 3,
    "cosine": 0.911427912597457
  },
  {
    "model": "text-embedding-3-small",
    "register": "memo",
    "chunk_limit": 1600,
    "config": "exact+balanced/weighted",
    "chunks": 3,
    "cosine": 0.9379545312022664
  },
  {
    "model": "text-embedding-3-small",
    "register": "memo",
    "chunk_limit": 1600,
    "config": "exact+balanced/unweighted",
    "chunks": 3,
    "cosine": 0.9379545312022664
  },
  {
    "model": "text-embedding-3-small",
    "register": "memo",
    "chunk_limit": 3200,
    "config": "exact+greedy/weighted",
    "chunks": 2,
    "cosine": 0.9888139109965072
  },
  {
    "model": "text-embedding-3-small",
    "register": "memo",
    "chunk_limit": 3200,
    "config": "exact+greedy/unweighted",
    "chunks": 2,
    "cosine": 0.8564703655992303
  },
  {
    "model": "text-embedding-3-small",
    "register": "memo",
    "chunk_limit": 3200,
    "config": "exact+balanced/weighted",
    "chunks": 2,
    "cosine": 0.9694625631518298
  },
  {
    "model": "text-embedding-3-small",
    "register": "memo",
    "chunk_limit": 3200,
    "config": "exact+balanced/unweighted",
    "chunks": 2,
    "cosine": 0.9694625631518298
  },
  {
    "model": "text-embedding-3-small",
    "register": "review",
    "chunk_limit": 800,
    "config": "exact+greedy/weighted",
    "chunks": 4,
    "cosine": 0.8943376644898445
  },
  {
    "model": "text-embedding-3-small",
    "register": "review",
    "chunk_limit": 800,
    "config": "exact+greedy/unweighted",
    "chunks": 4,
    "cosine": 0.8937838374846891
  },
  {
    "model": "text-embedding-3-small",
    "register": "review",
    "chunk_limit": 800,
    "config": "exact+balanced/weighted",
    "chunks": 4,
    "cosine": 0.9003077111303335
  },
  {
    "model": "text-embedding-3-small",
    "register": "review",
    "chunk_limit": 800,
    "config": "exact+balanced/unweighted",
    "chunks": 4,
    "cosine": 0.9002997207769681
  },
  {
    "model": "text-embedding-3-small",
    "register": "review",
    "chunk_limit": 1600,
    "config": "exact+greedy/weighted",
    "chunks": 2,
    "cosine": 0.9693488619549587
  },
  {
    "model": "text-embedding-3-small",
    "register": "review",
    "chunk_limit": 1600,
    "config": "exact+greedy/unweighted",
    "chunks": 2,
    "cosine": 0.9685939773868549
  },
  {
    "model": "text-embedding-3-small",
    "register": "review",
    "chunk_limit": 1600,
    "config": "exact+balanced/weighted",
    "chunks": 2,
    "cosine": 0.9661807444273959
  },
  {
    "model": "text-embedding-3-small",
    "register": "review",
    "chunk_limit": 1600,
    "config": "exact+balanced/unweighted",
    "chunks": 2,
    "cosine": 0.9661807444273959
  },
  {
    "model": "text-embedding-3-small",
    "register": "encyclopedia",
    "chunk_limit": 800,
    "config": "exact+greedy/weighted",
    "chunks": 5,
    "cosine": 0.9392862640357529
  },
  {
    "model": "text-embedding-3-small",
    "register": "encyclopedia",
    "chunk_limit": 800,
    "config": "exact+greedy/unweighted",
    "chunks": 5,
    "cosine": 0.9388041629072712
  },
  {
    "model": "text-embedding-3-small",
    "register": "encyclopedia",
    "chunk_limit": 800,
    "config": "exact+balanced/weighted",
    "chunks": 5,
    "cosine": 0.9242696383090272
  },
  {
    "model": "text-embedding-3-small",
    "register": "encyclopedia",
    "chunk_limit": 800,
    "config": "exact+balanced/unweighted",
    "chunks": 5,
    "cosine": 0.9242696383090272
  },
  {
    "model": "text-embedding-3-small",
    "register": "encyclopedia",
    "chunk_limit": 1600,
    "config": "exact+greedy/weighted",
    "chunks": 3,
    "cosine": 0.9691300498384388
  },
  {
    "model": "text-embedding-3-small",
    "register": "encyclopedia",
    "chunk_limit": 1600,
    "config": "exact+greedy/unweighted",
    "chunks": 3,
    "cosine": 0.9530099061375295
  },
  {
    "model": "text-embedding-3-small",
    "register": "encyclopedia",
    "chunk_limit": 1600,
    "config": "exact+balanced/weighted",
    "chunks": 3,
    "cosine": 0.9559655645621391
  },
  {
    "model": "text-embedding-3-small",
    "register": "encyclopedia",
    "chunk_limit": 1600,
    "config": "exact+balanced/unweighted",
    "chunks": 3,
    "cosine": 0.9559876824370069
  },
  {
    "model": "text-embedding-3-small",
    "register": "encyclopedia",
    "chunk_limit": 3200,
    "config": "exact+greedy/weighted",
    "chunks": 2,
    "cosine": 0.9914898834005675
  },
  {
    "model": "text-embedding-3-small",
    "register": "encyclopedia",
    "chunk_limit": 3200,
    "config": "exact+greedy/unweighted",
    "chunks": 2,
    "cosine": 0.932379948142258
  },
  {
    "model": "text-embedding-3-small",
    "register": "encyclopedia",
    "chunk_limit": 3200,
    "config": "exact+balanced/weighted",
    "chunks": 2,
    "cosine": 0.9618311518836582
  },
  {
    "model": "text-embedding-3-small",
    "register": "encyclopedia",
    "chunk_limit": 3200,
    "config": "exact+balanced/unweighted",
    "chunks": 2,
    "cosine": 0.9618311518836582
  },
  {
    "model": "text-embedding-3-large",
    "register": "fiction",
    "chunk_limit": 800,
    "config": "exact+greedy/weighted",
    "chunks": 7,
    "cosine": 0.8725742863547817
  },
  {
    "model": "text-embedding-3-large",
    "register": "fiction",
    "chunk_limit": 800,
    "config": "exact+greedy/unweighted",
    "chunks": 7,
    "cosine": 0.8526747402464222
  },
  {
    "model": "text-embedding-3-large",
    "register": "fiction",
    "chunk_limit": 800,
    "config": "exact+balanced/weighted",
    "chunks": 7,
    "cosine": 0.8682655697541166
  },
  {
    "model": "text-embedding-3-large",
    "register": "fiction",
    "chunk_limit": 800,
    "config": "exact+balanced/unweighted",
    "chunks": 7,
    "cosine": 0.8683502978144999
  },
  {
    "model": "text-embedding-3-large",
    "register": "fiction",
    "chunk_limit": 1600,
    "config": "exact+greedy/weighted",
    "chunks": 4,
    "cosine": 0.9293981307759025
  },
  {
    "model": "text-embedding-3-large",
    "register": "fiction",
    "chunk_limit": 1600,
    "config": "exact+greedy/unweighted",
    "chunks": 4,
    "cosine": 0.8820317475770435
  },
  {
    "model": "text-embedding-3-large",
    "register": "fiction",
    "chunk_limit": 1600,
    "config": "exact+balanced/weighted",
    "chunks": 4,
    "cosine": 0.9168991859285924
  },
  {
    "model": "text-embedding-3-large",
    "register": "fiction",
    "chunk_limit": 1600,
    "config": "exact+balanced/unweighted",
    "chunks": 4,
    "cosine": 0.9169192313366726
  },
  {
    "model": "text-embedding-3-large",
    "register": "fiction",
    "chunk_limit": 3200,
    "config": "exact+greedy/weighted",
    "chunks": 2,
    "cosine": 0.9521742437875029
  },
  {
    "model": "text-embedding-3-large",
    "register": "fiction",
    "chunk_limit": 3200,
    "config": "exact+greedy/unweighted",
    "chunks": 2,
    "cosine": 0.9505388735163692
  },
  {
    "model": "text-embedding-3-large",
    "register": "fiction",
    "chunk_limit": 3200,
    "config": "exact+balanced/weighted",
    "chunks": 2,
    "cosine": 0.9563534616441991
  },
  {
    "model": "text-embedding-3-large",
    "register": "fiction",
    "chunk_limit": 3200,
    "config": "exact+balanced/unweighted",
    "chunks": 2,
    "cosine": 0.9563534616441991
  },
  {
    "model": "text-embedding-3-large",
    "register": "technical",
    "chunk_limit": 800,
    "config": "exact+greedy/weighted",
    "chunks": 7,
    "cosine": 0.8692957933692302
  },
  {
    "model": "text-embedding-3-large",
    "register": "technical",
    "chunk_limit": 800,
    "config": "exact+greedy/unweighted",
    "chunks": 7,
    "cosine": 0.8637868287194482
  },
  {
    "model": "text-embedding-3-large",
    "register": "technical",
    "chunk_limit": 800,
    "config": "exact+balanced/weighted",
    "chunks": 7,
    "cosine": 0.8535008441835695
  },
  {
    "model": "text-embedding-3-large",
    "register": "technical",
    "chunk_limit": 800,
    "config": "exact+balanced/unweighted",
    "chunks": 7,
    "cosine": 0.8535113895111908
  },
  {
    "model": "text-embedding-3-large",
    "register": "technical",
    "chunk_limit": 1600,
    "config": "exact+greedy/weighted",
    "chunks": 4,
    "cosine": 0.9225805995794008
  },
  {
    "model": "text-embedding-3-large",
    "register": "technical",
    "chunk_limit": 1600,
    "config": "exact+greedy/unweighted",
    "chunks": 4,
    "cosine": 0.9002158818649625
  },
  {
    "model": "text-embedding-3-large",
    "register": "technical",
    "chunk_limit": 1600,
    "config": "exact+balanced/weighted",
    "chunks": 4,
    "cosine": 0.8772760109398264
  },
  {
    "model": "text-embedding-3-large",
    "register": "technical",
    "chunk_limit": 1600,
    "config": "exact+balanced/unweighted",
    "chunks": 4,
    "cosine": 0.8772684924952502
  },
  {
    "model": "text-embedding-3-large",
    "register": "technical",
    "chunk_limit": 3200,
    "config": "exact+greedy/weighted",
    "chunks": 2,
    "cosine": 0.961142685906412
  },
  {
    "model": "text-embedding-3-large",
    "register": "technical",
    "chunk_limit": 3200,
    "config": "exact+greedy/unweighted",
    "chunks": 2,
    "cosine": 0.9395961595954896
  },
  {
    "model": "text-embedding-3-large",
    "register": "technical",
    "chunk_limit": 3200,
    "config": "exact+balanced/weighted",
    "chunks": 2,
    "cosine": 0.9428777607385924
  },
  {
    "model": "text-embedding-3-large",
    "register": "technical",
    "chunk_limit": 3200,
    "config": "exact+balanced/unweighted",
    "chunks": 2,
    "cosine": 0.942859509264055
  },
  {
    "model": "text-embedding-3-large",
    "register": "chat",
    "chunk_limit": 800,
    "config": "exact+greedy/weighted",
    "chunks": 5,
    "cosine": 0.8914046117859982
  },
  {
    "model": "text-embedding-3-large",
    "register": "chat",
    "chunk_limit": 800,
    "config": "exact+greedy/unweighted",
    "chunks": 5,
    "cosine": 0.9015165812704059
  },
  {
    "model": "text-embedding-3-large",
    "register": "chat",
    "chunk_limit": 800,
    "config": "exact+balanced/weighted",
    "chunks": 5,
    "cosine": 0.888958955116375
  },
  {
    "model": "text-embedding-3-large",
    "register": "chat",
    "chunk_limit": 800,
    "config": "exact+balanced/unweighted",
    "chunks": 5,
    "cosine": 0.8890120142230198
  },
  {
    "model": "text-embedding-3-large",
    "register": "chat",
    "chunk_limit": 1600,
    "config": "exact+greedy/weighted",
    "chunks": 3,
    "cosine": 0.9109433157358814
  },
  {
    "model": "text-embedding-3-large",
    "register": "chat",
    "chunk_limit": 1600,
    "config": "exact+greedy/unweighted",
    "chunks": 3,
    "cosine": 0.9323571664926563
  },
  {
    "model": "text-embedding-3-large",
    "register": "chat",
    "chunk_limit": 1600,
    "config": "exact+balanced/weighted",
    "chunks": 3,
    "cosine": 0.9265006648352556
  },
  {
    "model": "text-embedding-3-large",
    "register": "chat",
    "chunk_limit": 1600,
    "config": "exact+balanced/unweighted",
    "chunks": 3,
    "cosine": 0.9265205986285565
  },
  {
    "model": "text-embedding-3-large",
    "register": "chat",
    "chunk_limit": 3200,
    "config": "exact+greedy/weighted",
    "chunks": 2,
    "cosine": 0.9328029108854164
  },
  {
    "model": "text-embedding-3-large",
    "register": "chat",
    "chunk_limit": 3200,
    "config": "exact+greedy/unweighted",
    "chunks": 2,
    "cosine": 0.9446556905449254
  },
  {
    "model": "text-embedding-3-large",
    "register": "chat",
    "chunk_limit": 3200,
    "config": "exact+balanced/weighted",
    "chunks": 2,
    "cosine": 0.9516333959219625
  },
  {
    "model": "text-embedding-3-large",
    "register": "chat",
    "chunk_limit": 3200,
    "config": "exact+balanced/unweighted",
    "chunks": 2,
    "cosine": 0.9516333959219625
  },
  {
    "model": "text-embedding-3-large",
    "register": "news",
    "chunk_limit": 800,
    "config": "exact+greedy/weighted",
    "chunks": 7,
    "cosine": 0.9246468656131022
  },
  {
    "model": "text-embedding-3-large",
    "register": "news",
    "chunk_limit": 800,
    "config": "exact+greedy/unweighted",
    "chunks": 7,
    "cosine": 0.9123957923236939
  },
  {
    "model": "text-embedding-3-large",
    "register": "news",
    "chunk_limit": 800,
    "config": "exact+balanced/weighted",
    "chunks": 7,
    "cosine": 0.9179468316654169
  },
  {
    "model": "text-embedding-3-large",
    "register": "news",
    "chunk_limit": 800,
    "config": "exact+balanced/unweighted",
    "chunks": 7,
    "cosine": 0.9179728377602997
  },
  {
    "model": "text-embedding-3-large",
    "register": "news",
    "chunk_limit": 1600,
    "config": "exact+greedy/weighted",
    "chunks": 4,
    "cosine": 0.958589272273143
  },
  {
    "model": "text-embedding-3-large",
    "register": "news",
    "chunk_limit": 1600,
    "config": "exact+greedy/unweighted",
    "chunks": 4,
    "cosine": 0.9159792571477957
  },
  {
    "model": "text-embedding-3-large",
    "register": "news",
    "chunk_limit": 1600,
    "config": "exact+balanced/weighted",
    "chunks": 4,
    "cosine": 0.951514100450437
  },
  {
    "model": "text-embedding-3-large",
    "register": "news",
    "chunk_limit": 1600,
    "config": "exact+balanced/unweighted",
    "chunks": 4,
    "cosine": 0.9515185244963626
  },
  {
    "model": "text-embedding-3-large",
    "register": "news",
    "chunk_limit": 3200,
    "config": "exact+greedy/weighted",
    "chunks": 2,
    "cosine": 0.9783034833033829
  },
  {
    "model": "text-embedding-3-large",
    "register": "news",
    "chunk_limit": 3200,
    "config": "exact+greedy/unweighted",
    "chunks": 2,
    "cosine": 0.9719360296544789
  },
  {
    "model": "text-embedding-3-large",
    "register": "news",
    "chunk_limit": 3200,
    "config": "exact+balanced/weighted",
    "chunks": 2,
    "cosine": 0.973007552101658
  },
  {
    "model": "text-embedding-3-large",
    "register": "news",
    "chunk_limit": 3200,
    "config": "exact+balanced/unweighted",
    "chunks": 2,
    "cosine": 0.9730056816979616
  },
  {
    "model": "text-embedding-3-large",
    "register": "legal",
    "chunk_limit": 800,
    "config": "exact+greedy/weighted",
    "chunks": 8,
    "cosine": 0.8762290161018007
  },
  {
    "model": "text-embedding-3-large",
    "register": "legal",
    "chunk_limit": 800,
    "config": "exact+greedy/unweighted",
    "chunks": 8,
    "cosine": 0.8618921965574035
  },
  {
    "model": "text-embedding-3-large",
    "register": "legal",
    "chunk_limit": 800,
    "config": "exact+balanced/weighted",
    "chunks": 8,
    "cosine": 0.8621933690613822
  },
  {
    "model": "text-embedding-3-large",
    "register": "legal",
    "chunk_limit": 800,
    "config": "exact+balanced/unweighted",
    "chunks": 8,
    "cosine": 0.8622122606337336
  },
  {
    "model": "text-embedding-3-large",
    "register": "legal",
    "chunk_limit": 1600,
    "config": "exact+greedy/weighted",
    "chunks": 4,
    "cosine": 0.9179491358333217
  },
  {
    "model": "text-embedding-3-large",
    "register": "legal",
    "chunk_limit": 1600,
    "config": "exact+greedy/unweighted",
    "chunks": 4,
    "cosine": 0.9152509319465725
  },
  {
    "model": "text-embedding-3-large",
    "register": "legal",
    "chunk_limit": 1600,
    "config": "exact+balanced/weighted",
    "chunks": 4,
    "cosine": 0.9163207092661529
  },
  {
    "model": "text-embedding-3-large",
    "register": "legal",
    "chunk_limit": 1600,
    "config": "exact+balanced/unweighted",
    "chunks": 4,
    "cosine": 0.9163195415885393
  },
  {
    "model": "text-embedding-3-large",
    "register": "legal",
    "chunk_limit": 3200,
    "config": "exact+greedy/weighted",
    "chunks": 2,
    "cosine": 0.9673150129260517
  },
  {
    "model": "text-embedding-3-large",
    "register": "legal",
    "chunk_limit": 3200,
    "config": "exact+greedy/unweighted",
    "chunks": 2,
    "cosine": 0.9594861490622085
  },
  {
    "model": "text-embedding-3-large",
    "register": "legal",
    "chunk_limit": 3200,
    "config": "exact+balanced/weighted",
    "chunks": 2,
    "cosine": 0.9588008218836657
  },
  {
    "model": "text-embedding-3-large",
    "register": "legal",
    "chunk_limit": 3200,
    "config": "exact+balanced/unweighted",
    "chunks": 2,
    "cosine": 0.9588008218836657
  },
  {
    "model": "text-embedding-3-large",
    "register": "academic",
    "chunk_limit": 800,
    "config": "exact+greedy/weighted",
    "chunks": 5,
    "cosine": 0.9089618687301707
  },
  {
    "model": "text-embedding-3-large",
    "register": "academic",
    "chunk_limit": 800,
    "config": "exact+greedy/unweighted",
    "chunks": 5,
    "cosine": 0.9083512086667455
  },
  {
    "model": "text-embedding-3-large",
    "register": "academic",
    "chunk_limit": 800,
    "config": "exact+balanced/weighted",
    "chunks": 5,
    "cosine": 0.9093429718060145
  },
  {
    "model": "text-embedding-3-large",
    "register": "academic",
    "chunk_limit": 800,
    "config": "exact+balanced/unweighted",
    "chunks": 5,
    "cosine": 0.9093295034679199
  },
  {
    "model": "text-embedding-3-large",
    "register": "academic",
    "chunk_limit": 1600,
    "config": "exact+greedy/weighted",
    "chunks": 3,
    "cosine": 0.9493023947420914
  },
  {
    "model": "text-embedding-3-large",
    "register": "academic",
    "chunk_limit": 1600,
    "config": "exact+greedy/unweighted",
    "chunks": 3,
    "cosine": 0.9301697587923871
  },
  {
    "model": "text-embedding-3-large",
    "register": "academic",
    "chunk_limit": 1600,
    "config": "exact+balanced/weighted",
    "chunks": 3,
    "cosine": 0.934960064363141
  },
  {
    "model": "text-embedding-3-large",
    "register": "academic",
    "chunk_limit": 1600,
    "config": "exact+balanced/unweighted",
    "chunks": 3,
    "cosine": 0.9349217876611146
  },
  {
    "model": "text-embedding-3-large",
    "register": "academic",
    "chunk_limit": 3200,
    "config": "exact+greedy/weighted",
    "chunks": 2,
    "cosine": 0.9877894160395762
  },
  {
    "model": "text-embedding-3-large",
    "register": "academic",
    "chunk_limit": 3200,
    "config": "exact+greedy/unweighted",
    "chunks": 2,
    "cosine": 0.9206149072731512
  },
  {
    "model": "text-embedding-3-large",
    "register": "academic",
    "chunk_limit": 3200,
    "config": "exact+balanced/weighted",
    "chunks": 2,
    "cosine": 0.9583997184028397
  },
  {
    "model": "text-embedding-3-large",
    "register": "academic",
    "chunk_limit": 3200,
    "config": "exact+balanced/unweighted",
    "chunks": 2,
    "cosine": 0.9583997184028397
  },
  {
    "model": "text-embedding-3-large",
    "register": "recipe",
    "chunk_limit": 800,
    "config": "exact+greedy/weighted",
    "chunks": 4,
    "cosine": 0.9085917752009645
  },
  {
    "model": "text-embedding-3-large",
    "register": "recipe",
    "chunk_limit": 800,
    "config": "exact+greedy/unweighted",
    "chunks": 4,
    "cosine": 0.9098349744781539
  },
  {
    "model": "text-embedding-3-large",
    "register": "recipe",
    "chunk_limit": 800,
    "config": "exact+balanced/weighted",
    "chunks": 4,
    "cosine": 0.9121437574370774
  },
  {
    "model": "text-embedding-3-large",
    "register": "recipe",
    "chunk_limit": 800,
    "config": "exact+balanced/unweighted",
    "chunks": 4,
    "cosine": 0.9121498741143104
  },
  {
    "model": "text-embedding-3-large",
    "register": "recipe",
    "chunk_limit": 1600,
    "config": "exact+greedy/weighted",
    "chunks": 2,
    "cosine": 0.9555594753515898
  },
  {
    "model": "text-embedding-3-large",
    "register": "recipe",
    "chunk_limit": 1600,
    "config": "exact+greedy/unweighted",
    "chunks": 2,
    "cosine": 0.9542723567879481
  },
  {
    "model": "text-embedding-3-large",
    "register": "recipe",
    "chunk_limit": 1600,
    "config": "exact+balanced/weighted",
    "chunks": 2,
    "cosine": 0.9595537486349814
  },
  {
    "model": "text-embedding-3-large",
    "register": "recipe",
    "chunk_limit": 1600,
    "config": "exact+balanced/unweighted",
    "chunks": 2,
    "cosine": 0.9595410150516287
  },
  {
    "model": "text-embedding-3-large",
    "register": "memo",
    "chunk_limit": 800,
    "config": "exact+greedy/weighted",
    "chunks": 5,
    "cosine": 0.8852235402285263
  },
  {
    "model": "text-embedding-3-large",
    "register": "memo",
    "chunk_limit": 800,
    "config": "exact+greedy/unweighted",
    "chunks": 5,
    "cosine": 0.8805573545185522
  },
  {
    "model": "text-embedding-3-large",
    "register": "memo",
    "chunk_limit": 800,
    "config": "exact+balanced/weighted",
    "chunks": 5,
    "cosine": 0.8870884027265513
  },
  {
    "model": "text-embedding-3-large",
    "register": "memo",
    "chunk_limit": 800,
    "config": "exact+balanced/unweighted",
    "chunks": 5,
    "cosine": 0.8870611786430033
  },
  {
    "model": "text-embedding-3-large",
    "register": "memo",
    "chunk_limit": 1600,
    "config": "exact+greedy/weighted",
    "chunks": 3,
    "cosine": 0.935595937723273
  },
  {
    "model": "text-embedding-3-large",
    "register": "memo",
    "chunk_limit": 1600,
    "config": "exact+greedy/unweighted",
    "chunks": 3,
    "cosine": 0.8904612012785335
  },
  {
    "model": "text-embedding-3-large",
    "register": "memo",
    "chunk_limit": 1600,
    "config": "exact+balanced/weighted",
    "chunks": 3,
    "cosine": 0.9236257268194973
  },
  {
    "model": "text-embedding-3-large",
    "register": "memo",
    "chunk_limit": 1600,
    "config": "exact+balanced/unweighted",
    "chunks": 3,
    "cosine": 0.9236257268194973
  },
  {
    "model": "text-embedding-3-large",
    "register": "memo",
    "chunk_limit": 3200,
    "config": "exact+greedy/weighted",
    "chunks": 2,
    "cosine": 0.9888784682282963
  },
  {
    "model": "text-embedding-3-large",
    "register": "memo",
    "chunk_limit": 3200,
    "config": "exact+greedy/unweighted",
    "chunks": 2,
    "cosine": 0.8425945625327097
  },
  {
    "model": "text-embedding-3-large",
    "register": "memo",
    "chunk_limit": 3200,
    "config": "exact+balanced/weighted",
    "chunks": 2,
    "cosine": 0.9477363642341258
  },
  {
    "model": "text-embedding-3-large",
    "register": "memo",
    "chunk_limit": 3200,
    "config": "exact+balanced/unweighted",
    "chunks": 2,
    "cosine": 0.9477363642341258
  },
  {
    "model": "text-embedding-3-large",
    "register": "review",
    "chunk_limit": 800,
    "config": "exact+greedy/weighted",
    "chunks": 4,
    "cosine": 0.9155605920011464
  },
  {
    "model": "text-embedding-3-large",
    "register": "review",
    "chunk_limit": 800,
    "config": "exact+greedy/unweighted",
    "chunks": 4,
    "cosine": 0.9161114250863913
  },
  {
    "model": "text-embedding-3-large",
    "register": "review",
    "chunk_limit": 800,
    "config": "exact+balanced/weighted",
    "chunks": 4,
    "cosine": 0.9164096698012906
  },
  {
    "model": "text-embedding-3-large",
    "register": "review",
    "chunk_limit": 800,
    "config": "exact+balanced/unweighted",
    "chunks": 4,
    "cosine": 0.9164222476637314
  },
  {
    "model": "text-embedding-3-large",
    "register": "review",
    "chunk_limit": 1600,
    "config": "exact+greedy/weighted",
    "chunks": 2,
    "cosine": 0.9636697821940844
  },
  {
    "model": "text-embedding-3-large",
    "register": "review",
    "chunk_limit": 1600,
    "config": "exact+greedy/unweighted",
    "chunks": 2,
    "cosine": 0.9626390694884561
  },
  {
    "model": "text-embedding-3-large",
    "register": "review",
    "chunk_limit": 1600,
    "config": "exact+balanced/weighted",
    "chunks": 2,
    "cosine": 0.9621731533155191
  },
  {
    "model": "text-embedding-3-large",
    "register": "review",
    "chunk_limit": 1600,
    "config": "exact+balanced/unweighted",
    "chunks": 2,
    "cosine": 0.9621731533155191
  },
  {
    "model": "text-embedding-3-large",
    "register": "encyclopedia",
    "chunk_limit": 800,
    "config": "exact+greedy/weighted",
    "chunks": 5,
    "cosine": 0.9284587937641582
  },
  {
    "model": "text-embedding-3-large",
    "register": "encyclopedia",
    "chunk_limit": 800,
    "config": "exact+greedy/unweighted",
    "chunks": 5,
    "cosine": 0.9310569323372485
  },
  {
    "model": "text-embedding-3-large",
    "register": "encyclopedia",
    "chunk_limit": 800,
    "config": "exact+balanced/weighted",
    "chunks": 5,
    "cosine": 0.9070057119361145
  },
  {
    "model": "text-embedding-3-large",
    "register": "encyclopedia",
    "chunk_limit": 800,
    "config": "exact+balanced/unweighted",
    "chunks": 5,
    "cosine": 0.9070057119361145
  },
  {
    "model": "text-embedding-3-large",
    "register": "encyclopedia",
    "chunk_limit": 1600,
    "config": "exact+greedy/weighted",
    "chunks": 3,
    "cosine": 0.9667302361494651
  },
  {
    "model": "text-embedding-3-large",
    "register": "encyclopedia",
    "chunk_limit": 1600,
    "config": "exact+greedy/unweighted",
    "chunks": 3,
    "cosine": 0.9569990642677528
  },
  {
    "model": "text-embedding-3-large",
    "register": "encyclopedia",
    "chunk_limit": 1600,
    "config": "exact+balanced/weighted",
    "chunks": 3,
    "cosine": 0.9380974969643602
  },
  {
    "model": "text-embedding-3-large",
    "register": "encyclopedia",
    "chunk_limit": 1600,
    "config": "exact+balanced/unweighted",
    "chunks": 3,
    "cosine": 0.9381025413144787
  },
  {
    "model": "text-embedding-3-large",
    "register": "encyclopedia",
    "chunk_limit": 3200,
    "config": "exact+greedy/weighted",
    "chunks": 2,
    "cosine": 0.9920734926735856
  },
  {
    "model": "text-embedding-3-large",
    "register": "encyclopedia",
    "chunk_limit": 3200,
    "config": "exact+greedy/unweighted",
    "chunks": 2,
    "cosine": 0.9428675288329597
  },
  {
    "model": "text-embedding-3-large",
    "register": "encyclopedia",
    "chunk_limit": 3200,
    "config": "exact+balanced/weighted",
    "chunks": 2,
    "cosine": 0.9438180658727109
  },
  {
    "model": "text-embedding-3-large",
    "register": "encyclopedia",
    "chunk_limit": 3200,
    "config": "exact+balanced/unweighted",
    "chunks": 2,
    "cosine": 0.9438180658727109
  }
]

@edwinyyyu

Copy link
Copy Markdown
Contributor Author

Probe attachment 2/2 — special-token server behavior (kept out of the branch deliberately; attached for reviewers)

probe_special_token_sanitization.py
"""Probe: does the OpenAI embeddings endpoint still error on special tokens?

Context: OpenAIEmbedder._SPECIAL_TOKEN_PATTERN.sub("") is a single
non-recursive pass, so a nested occurrence like "<|endof<|endoftext|>text|>"
sanitizes to a live "<|endoftext|>" that is sent to the provider verbatim.
This probe (1) demonstrates the reconstruction locally and (2) sends raw and
reconstructed tokens to each OpenAI embedding model to record current
server-side behavior.

Run: uv run python probe_special_token_sanitization.py
"""

import json
import os
import re
import sys
from pathlib import Path

import openai

ENV_PATH = Path(
    "/Users/eyu/edwinyyyu/mmcc/segment_store/evaluation/event_memory/locomo/.env"
)

SPECIAL_TOKENS = [
    "<|endoftext|>",
    "<|im_start|>",
    "<|im_end|>",
    "<|fim_prefix|>",
    "<|fim_middle|>",
    "<|fim_suffix|>",
    "<|endofprompt|>",
]

# Mirrors OpenAIEmbedder._SPECIAL_TOKEN_PATTERN.
SPECIAL_TOKEN_PATTERN = re.compile(
    "|".join(re.escape(token) for token in SPECIAL_TOKENS)
)

MODELS = [
    "text-embedding-3-small",
    "text-embedding-3-large",
    "text-embedding-ada-002",
]


def load_api_key() -> str:
    for line in ENV_PATH.read_text().splitlines():
        line = line.strip()
        if line.startswith("OPENAI_API_KEY"):
            return line.split("=", 1)[1].strip().strip('"').strip("'")
    raise RuntimeError(f"OPENAI_API_KEY not found in {ENV_PATH}")


def sanitize(text: str) -> str:
    return SPECIAL_TOKEN_PATTERN.sub("", text) or "."


def nest(token: str) -> str:
    # Splice the full token inside itself: "<|endof" + token + "text|>"
    # for "<|endoftext|>". One-pass removal of the inner occurrence
    # reconstructs the outer token.
    midpoint = len(token) // 2
    return token[:midpoint] + token + token[midpoint:]


def try_embed(client: openai.OpenAI, model: str, text: str) -> dict:
    try:
        response = client.embeddings.create(input=text, model=model)
        return {
            "outcome": "OK",
            "detail": f"dims={len(response.data[0].embedding)}",
        }
    except openai.APIStatusError as err:
        body = err.body if isinstance(err.body, dict) else {}
        message = body.get("message") or str(err)
        return {
            "outcome": f"HTTP {err.status_code} {type(err).__name__}",
            "detail": str(message)[:200],
        }
    except openai.OpenAIError as err:
        return {"outcome": type(err).__name__, "detail": str(err)[:200]}


def main() -> None:
    client = openai.OpenAI(api_key=load_api_key(), max_retries=0)

    print("=== Local sanitizer check (no API) ===")
    for token in SPECIAL_TOKENS:
        nested = nest(token)
        sanitized = sanitize(nested)
        reconstructs = sanitized == token
        print(
            f"{token:18} nested={nested!r} -> sanitized={sanitized!r} "
            f"reconstructs_live_token={reconstructs}"
        )
        if not reconstructs:
            print("  UNEXPECTED: nesting did not reconstruct; check pattern")

    results = []
    for model in MODELS:
        print(f"\n=== {model} ===")
        control = try_embed(client, model, "hello world")
        print(f"{'control (clean text)':42} {control['outcome']}")
        results.append({"model": model, "case": "control", **control})

        for token in SPECIAL_TOKENS:
            for case_name, payload in [
                ("raw token alone", token),
                ("token embedded in text", f"hello {token} world"),
                ("sanitized nested input", sanitize(nest(token))),
            ]:
                result = try_embed(client, model, payload)
                label = f"{token} ({case_name})"
                print(f"{label:42} {result['outcome']}  {result['detail']}")
                results.append(
                    {
                        "model": model,
                        "case": case_name,
                        "token": token,
                        "payload": payload,
                        **result,
                    }
                )

    out_path = Path(__file__).with_suffix(".results.json")
    out_path.write_text(json.dumps(results, indent=2))
    print(f"\nWrote {len(results)} results to {out_path}")


if __name__ == "__main__":
    if not ENV_PATH.exists():
        sys.exit(f"Missing env file: {ENV_PATH}")
    os.environ.pop("OPENAI_API_KEY", None)
    main()
probe_special_token_sanitization.results.json
[
  {
    "model": "text-embedding-3-small",
    "case": "control",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-3-small",
    "case": "raw token alone",
    "token": "<|endoftext|>",
    "payload": "<|endoftext|>",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-3-small",
    "case": "token embedded in text",
    "token": "<|endoftext|>",
    "payload": "hello <|endoftext|> world",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-3-small",
    "case": "sanitized nested input",
    "token": "<|endoftext|>",
    "payload": "<|endoftext|>",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-3-small",
    "case": "raw token alone",
    "token": "<|im_start|>",
    "payload": "<|im_start|>",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-3-small",
    "case": "token embedded in text",
    "token": "<|im_start|>",
    "payload": "hello <|im_start|> world",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-3-small",
    "case": "sanitized nested input",
    "token": "<|im_start|>",
    "payload": "<|im_start|>",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-3-small",
    "case": "raw token alone",
    "token": "<|im_end|>",
    "payload": "<|im_end|>",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-3-small",
    "case": "token embedded in text",
    "token": "<|im_end|>",
    "payload": "hello <|im_end|> world",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-3-small",
    "case": "sanitized nested input",
    "token": "<|im_end|>",
    "payload": "<|im_end|>",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-3-small",
    "case": "raw token alone",
    "token": "<|fim_prefix|>",
    "payload": "<|fim_prefix|>",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-3-small",
    "case": "token embedded in text",
    "token": "<|fim_prefix|>",
    "payload": "hello <|fim_prefix|> world",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-3-small",
    "case": "sanitized nested input",
    "token": "<|fim_prefix|>",
    "payload": "<|fim_prefix|>",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-3-small",
    "case": "raw token alone",
    "token": "<|fim_middle|>",
    "payload": "<|fim_middle|>",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-3-small",
    "case": "token embedded in text",
    "token": "<|fim_middle|>",
    "payload": "hello <|fim_middle|> world",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-3-small",
    "case": "sanitized nested input",
    "token": "<|fim_middle|>",
    "payload": "<|fim_middle|>",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-3-small",
    "case": "raw token alone",
    "token": "<|fim_suffix|>",
    "payload": "<|fim_suffix|>",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-3-small",
    "case": "token embedded in text",
    "token": "<|fim_suffix|>",
    "payload": "hello <|fim_suffix|> world",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-3-small",
    "case": "sanitized nested input",
    "token": "<|fim_suffix|>",
    "payload": "<|fim_suffix|>",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-3-small",
    "case": "raw token alone",
    "token": "<|endofprompt|>",
    "payload": "<|endofprompt|>",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-3-small",
    "case": "token embedded in text",
    "token": "<|endofprompt|>",
    "payload": "hello <|endofprompt|> world",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-3-small",
    "case": "sanitized nested input",
    "token": "<|endofprompt|>",
    "payload": "<|endofprompt|>",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-3-large",
    "case": "control",
    "outcome": "OK",
    "detail": "dims=3072"
  },
  {
    "model": "text-embedding-3-large",
    "case": "raw token alone",
    "token": "<|endoftext|>",
    "payload": "<|endoftext|>",
    "outcome": "OK",
    "detail": "dims=3072"
  },
  {
    "model": "text-embedding-3-large",
    "case": "token embedded in text",
    "token": "<|endoftext|>",
    "payload": "hello <|endoftext|> world",
    "outcome": "OK",
    "detail": "dims=3072"
  },
  {
    "model": "text-embedding-3-large",
    "case": "sanitized nested input",
    "token": "<|endoftext|>",
    "payload": "<|endoftext|>",
    "outcome": "OK",
    "detail": "dims=3072"
  },
  {
    "model": "text-embedding-3-large",
    "case": "raw token alone",
    "token": "<|im_start|>",
    "payload": "<|im_start|>",
    "outcome": "OK",
    "detail": "dims=3072"
  },
  {
    "model": "text-embedding-3-large",
    "case": "token embedded in text",
    "token": "<|im_start|>",
    "payload": "hello <|im_start|> world",
    "outcome": "OK",
    "detail": "dims=3072"
  },
  {
    "model": "text-embedding-3-large",
    "case": "sanitized nested input",
    "token": "<|im_start|>",
    "payload": "<|im_start|>",
    "outcome": "OK",
    "detail": "dims=3072"
  },
  {
    "model": "text-embedding-3-large",
    "case": "raw token alone",
    "token": "<|im_end|>",
    "payload": "<|im_end|>",
    "outcome": "OK",
    "detail": "dims=3072"
  },
  {
    "model": "text-embedding-3-large",
    "case": "token embedded in text",
    "token": "<|im_end|>",
    "payload": "hello <|im_end|> world",
    "outcome": "OK",
    "detail": "dims=3072"
  },
  {
    "model": "text-embedding-3-large",
    "case": "sanitized nested input",
    "token": "<|im_end|>",
    "payload": "<|im_end|>",
    "outcome": "OK",
    "detail": "dims=3072"
  },
  {
    "model": "text-embedding-3-large",
    "case": "raw token alone",
    "token": "<|fim_prefix|>",
    "payload": "<|fim_prefix|>",
    "outcome": "OK",
    "detail": "dims=3072"
  },
  {
    "model": "text-embedding-3-large",
    "case": "token embedded in text",
    "token": "<|fim_prefix|>",
    "payload": "hello <|fim_prefix|> world",
    "outcome": "OK",
    "detail": "dims=3072"
  },
  {
    "model": "text-embedding-3-large",
    "case": "sanitized nested input",
    "token": "<|fim_prefix|>",
    "payload": "<|fim_prefix|>",
    "outcome": "OK",
    "detail": "dims=3072"
  },
  {
    "model": "text-embedding-3-large",
    "case": "raw token alone",
    "token": "<|fim_middle|>",
    "payload": "<|fim_middle|>",
    "outcome": "OK",
    "detail": "dims=3072"
  },
  {
    "model": "text-embedding-3-large",
    "case": "token embedded in text",
    "token": "<|fim_middle|>",
    "payload": "hello <|fim_middle|> world",
    "outcome": "OK",
    "detail": "dims=3072"
  },
  {
    "model": "text-embedding-3-large",
    "case": "sanitized nested input",
    "token": "<|fim_middle|>",
    "payload": "<|fim_middle|>",
    "outcome": "OK",
    "detail": "dims=3072"
  },
  {
    "model": "text-embedding-3-large",
    "case": "raw token alone",
    "token": "<|fim_suffix|>",
    "payload": "<|fim_suffix|>",
    "outcome": "OK",
    "detail": "dims=3072"
  },
  {
    "model": "text-embedding-3-large",
    "case": "token embedded in text",
    "token": "<|fim_suffix|>",
    "payload": "hello <|fim_suffix|> world",
    "outcome": "OK",
    "detail": "dims=3072"
  },
  {
    "model": "text-embedding-3-large",
    "case": "sanitized nested input",
    "token": "<|fim_suffix|>",
    "payload": "<|fim_suffix|>",
    "outcome": "OK",
    "detail": "dims=3072"
  },
  {
    "model": "text-embedding-3-large",
    "case": "raw token alone",
    "token": "<|endofprompt|>",
    "payload": "<|endofprompt|>",
    "outcome": "OK",
    "detail": "dims=3072"
  },
  {
    "model": "text-embedding-3-large",
    "case": "token embedded in text",
    "token": "<|endofprompt|>",
    "payload": "hello <|endofprompt|> world",
    "outcome": "OK",
    "detail": "dims=3072"
  },
  {
    "model": "text-embedding-3-large",
    "case": "sanitized nested input",
    "token": "<|endofprompt|>",
    "payload": "<|endofprompt|>",
    "outcome": "OK",
    "detail": "dims=3072"
  },
  {
    "model": "text-embedding-ada-002",
    "case": "control",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-ada-002",
    "case": "raw token alone",
    "token": "<|endoftext|>",
    "payload": "<|endoftext|>",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-ada-002",
    "case": "token embedded in text",
    "token": "<|endoftext|>",
    "payload": "hello <|endoftext|> world",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-ada-002",
    "case": "sanitized nested input",
    "token": "<|endoftext|>",
    "payload": "<|endoftext|>",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-ada-002",
    "case": "raw token alone",
    "token": "<|im_start|>",
    "payload": "<|im_start|>",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-ada-002",
    "case": "token embedded in text",
    "token": "<|im_start|>",
    "payload": "hello <|im_start|> world",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-ada-002",
    "case": "sanitized nested input",
    "token": "<|im_start|>",
    "payload": "<|im_start|>",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-ada-002",
    "case": "raw token alone",
    "token": "<|im_end|>",
    "payload": "<|im_end|>",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-ada-002",
    "case": "token embedded in text",
    "token": "<|im_end|>",
    "payload": "hello <|im_end|> world",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-ada-002",
    "case": "sanitized nested input",
    "token": "<|im_end|>",
    "payload": "<|im_end|>",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-ada-002",
    "case": "raw token alone",
    "token": "<|fim_prefix|>",
    "payload": "<|fim_prefix|>",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-ada-002",
    "case": "token embedded in text",
    "token": "<|fim_prefix|>",
    "payload": "hello <|fim_prefix|> world",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-ada-002",
    "case": "sanitized nested input",
    "token": "<|fim_prefix|>",
    "payload": "<|fim_prefix|>",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-ada-002",
    "case": "raw token alone",
    "token": "<|fim_middle|>",
    "payload": "<|fim_middle|>",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-ada-002",
    "case": "token embedded in text",
    "token": "<|fim_middle|>",
    "payload": "hello <|fim_middle|> world",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-ada-002",
    "case": "sanitized nested input",
    "token": "<|fim_middle|>",
    "payload": "<|fim_middle|>",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-ada-002",
    "case": "raw token alone",
    "token": "<|fim_suffix|>",
    "payload": "<|fim_suffix|>",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-ada-002",
    "case": "token embedded in text",
    "token": "<|fim_suffix|>",
    "payload": "hello <|fim_suffix|> world",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-ada-002",
    "case": "sanitized nested input",
    "token": "<|fim_suffix|>",
    "payload": "<|fim_suffix|>",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-ada-002",
    "case": "raw token alone",
    "token": "<|endofprompt|>",
    "payload": "<|endofprompt|>",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-ada-002",
    "case": "token embedded in text",
    "token": "<|endofprompt|>",
    "payload": "hello <|endofprompt|> world",
    "outcome": "OK",
    "detail": "dims=1536"
  },
  {
    "model": "text-embedding-ada-002",
    "case": "sanitized nested input",
    "token": "<|endofprompt|>",
    "payload": "<|endofprompt|>",
    "outcome": "OK",
    "detail": "dims=1536"
  }
]

@github-actions

Copy link
Copy Markdown
Contributor

This pull request has been automatically marked as stale because it has not had recent activity. It will be closed in 14 days if no further activity occurs. If you are still working on this, please push a commit or leave a comment. Reviewers: please respond, or add the keep-open label if this PR should be held open for a longer review cycle.

@github-actions github-actions Bot added the Stale label Jul 27, 2026
@edwinyyyu edwinyyyu added the keep-open Prevents the auto-close task from closing this issue. label Jul 27, 2026
@edwinyyyu
edwinyyyu force-pushed the embedding_chunking_fidelity branch 2 times, most recently from 1d939f9 to b1b5c73 Compare August 25, 2026 23:13
@edwinyyyu edwinyyyu mentioned this pull request Sep 8, 2026
13 of 20 tasks
@edwinyyyu
edwinyyyu force-pushed the embedding_chunking_fidelity branch from b1b5c73 to 896ddbb Compare September 16, 2026 23:41
@edwinyyyu edwinyyyu changed the title Improve embedding fidelity [speedkick port 5/17] Improve embedding fidelity (port of #1587) Sep 16, 2026
This was referenced Sep 16, 2026
Improve embedding fidelity

Signed-off-by: Edwin Yu <[email protected]>
@edwinyyyu
edwinyyyu force-pushed the embedding_chunking_fidelity branch from 896ddbb to 36a4595 Compare September 17, 2026 00:01

@marvinyu-memverge marvinyu-memverge left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed at d2004376. The measurement is the best part of this PR and I'd take the change on its strength -- greedy+weighted wins on all three encoders, and including the two structural checks (greedy+unweighted ranking last, balanced being weighting-indifferent to five decimals) is what makes it readable as signal rather than a lucky corpus. Two asks, both about coverage of the evidence rather than the change being wrong.

1. The special-token probe covers api.openai.com; the sanitizer it removes covered every backend this class serves.

The probe is thorough for what it tests -- 7 tokens x 3 cases x 3 models, all OK, so "OpenAI no longer 500s" is well established. But OpenAIEmbedder is pointed at non-OpenAI endpoints as a shipped, documented configuration: sample_configs/episodic_memory_config.cpu.sample has two of them under provider: openai -- ollama_embedder (nomic-embed-text) and openai_compatible_embedder (text-embedding-v4) -- and OpenAIEmbedderConf.model is described as "OpenAI Embeddings API-compatible model". Those tokenizers aren't OpenAI's and weren't probed. Since MemMachine stores chat transcripts, text containing a literal <|endoftext|> is more likely here than in most corpora, and the failure mode is a provider 4xx/5xx that exhausts retries into ExternalServiceAPIError -- an ingest failure that looks random because it's data-dependent.

Not asking you to probe Ollama and DashScope. Either is fine by me: keep the sanitizer and scope it to non-OpenAI base_url, or drop it as you have and say in the PR body that OpenAI-compatible backends are out of the probe's scope, so the next person who hits it knows where to look.

2. Neither half of the change is pinned by a test.

test_embed_oversized_input_with_no_max_input_length asserts call_count >= 2, which passes under balanced and greedy alike -- the diff updates its comment but not its assertion, so nothing in the suite distinguishes the old policy from the new. Your own grid says greedy+unweighted is the worst of the four configurations, so a later edit that reverts just the weighting lands in the worst cell and no test fails; the damage is cosine drift, with no error to notice. A small unit test with two chunks of very different length, asserting the merged vector is the length-weighted average and not the plain mean, would pin the half that's easiest to lose.

Observation, not an ask: max_input_length has no upper bound, and cluster_texts is still called with max_total_input_length_per_request (75,000) while chunking now uses effective_max. Above 75,000 that's the #1298 crash again, and greedy widens the window balanced used to cover: at max_input_length=80_000 a 100,000-char input gives balanced [50000, 50000] (fine) and greedy [80000, 20000] (ValueError -> 500). The default path is safe and your regression test still holds -- every shipped sample sets 2048, so I don't think anything reaches this today. Flagging it because the fix is one line (min(self._max_input_length or MAX, MAX)) if you want to close it while you're in here.

Verified while reading: the zero-chunk case is closed on all three embedders (chunk_text("") returns [], but OpenAI and Bedrock coerce or "." first and SentenceTransformer guards with and input_text, and max(len(chunk), 1) floors every weight); zip(..., strict=True) can't mismatch because unflatten_like already raises on a shape mismatch; the sweep is complete, with no embedder left on balanced+unweighted, so merged embeddings don't diverge by backend; and the branch is 0 behind main, so the #1646 lock refresh that touched these modules is absorbed.

malatewang and others added 4 commits September 18, 2026 09:50
A configured max_input_length above max_total_input_length_per_request
(75,000 code points) can never be honored: no chunk larger than the cap
fits in any request. Before this change such a limit reached cluster_texts
with an oversized chunk and raised ValueError (HTTP 500), the MemMachine#1298 failure
for that configuration. Greedy chunking widened which input lengths hit it:
at max_input_length=80_000, balanced chunking crashed on 76,000-80,000 and
160,000-char inputs; greedy crashed on everything above 75,000.

The regression test configures max_input_length=80_000 with a 100,000-char
input; it fails before the clamp with the cluster_texts ValueError and
passes after with two requests (75,000 + 25,000 chars).

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
@edwinyyyu

Copy link
Copy Markdown
Contributor Author

Written by Claude Code from Edwin's positions. Paragraphs marked (Edwin) are his; those marked (Claude) are measurements or additions of mine that he has not separately asserted.

1. Sanitizer

(Edwin) It was a bug on OpenAI's servers, 500s on inputs containing these tokens, and the sanitizer was a temporary workaround for it. Any worthwhile embedding endpoint handles special tokens in input text correctly; OpenAI got special treatment because of its size. If a third-party endpoint is chosen and fails on special tokens, the remedy is to stop using that endpoint, not for MemMachine to pre-edit input for every provider that might mis-tokenize.

(Claude) The removed comment in the source said exactly that ("Tokens that cause 500 errors on text-embedding-3-small"), and the attached probe shows the bug is gone: all 66 cases (3 models x {control, 7 tokens x 3 input forms}) return 200. The sanitizer also wasn't free: it deleted the literal token from the text before embedding, so a stored transcript that quotes <|endoftext|> was embedded as if it didn't; the "MemMachine stores chat transcripts" argument cuts the other way. The PR body now carries this reasoning and states that OpenAI-compatible endpoints are outside the probe's scope.

2. Tests

(Edwin) A test asserting "merged is the length-weighted average, not the plain mean" pins the code, not a behavioral contract. Any future change to the code requires the same change to the test, for the same reason, so the code and the test are bijective and the test catches nothing. There is also no point pinning one specific way to chunk text for embedding when a better approach may come later.

(Claude) The contract (closeness to the whole-text embedding) is only observable against a real encoder. I checked whether the encoder the CI integration job already loads (all-MiniLM-L6-v2) could carry a test that fails before this change and passes after: it can't. At its 256-token window the before and after cells tie (balanced+unweighted 0.9222 vs greedy+weighted 0.9213 on a 10-text corpus, 10/30 wins; +0.007 at t=1.7 on a second corpus), though greedy+unweighted is clearly worst there too (0.9008). The measurement in the PR body is the evidence for the change; the body now says the fidelity change is not unit-pinned and why.

3. Observation

(Edwin) Why is this an observation rather than an ask? Taking the fix since it's simple, with a test that fails before the change and passes after.

(Claude) Clamped in 7dae4a2: effective_max = min(max_input_length or cap, cap). Greedy did widen the window: at max_input_length=80_000, balanced chunking crashed on 76,000-80,000 and 160,000-char inputs; greedy crashed on everything above 75,000. The regression test configures max_input_length=80_000 with a 100,000-char input; against the unclamped source it fails with the cluster_texts ValueError, and after the clamp it passes with two requests (75,000 + 25,000 chars). PR body updated with the clamp.

@marvinyu-memverge

Copy link
Copy Markdown
Collaborator

Merged, so nothing here needs action - answering the question and closing out my end.

On the observation. Reachability is the whole of it, and I read it off the producer rather than the threat model. max_input_length defaults to None, which resolves to the 75,000 cap, and every shipped sample sets it explicitly: sample_configs/episodic_memory_config.cpu.sample and .gpu.sample, 2048 at all eight sites, 36x below the trigger. So the mechanism was real and reachable by some configuration, but nothing we ship reaches it - which is the observation cell rather than the ask cell. Taking it anyway was yours to call and I think the right one: the clamp is three lines and the test genuinely discriminates, which is a better trade than carrying a known 500 for whoever first raises that limit.

On the sanitizer. Your point that it deleted the literal token from the text before embedding is a good one and it cuts against the argument I made. I used "MemMachine stores chat transcripts" as the reachability path for a mis-tokenizing endpoint, without weighing that the workaround was silently corrupting those same transcripts on the endpoint it did cover. Stating in the body that OpenAI-compatible endpoints are outside the probe's scope is the right resolution - it leaves the gap visible to whoever points the config at one.

On the tests. The MiniLM check answers the ask rather than deflecting it. I asked for the change to be pinned; showing that the encoder CI already loads cannot produce a test that fails before and passes after is the fact I was missing, and the body saying the fidelity change is not unit-pinned and why is what I actually wanted out of it.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

keep-open Prevents the auto-close task from closing this issue. Stale

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants