Skip to content

Make export friendlier to SQL - #11

Merged
lhotari merged 1 commit into
mainfrom
export-for-sql
Sep 23, 2026
Merged

lhotari merged 1 commit into
mainfrom
export-for-sql

Conversation

@lhotari

@lhotari lhotari commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

Second PR of the 0.5.0 correlator stack (GitHub stack #15; spec: export-for-sql.md). Based on #10 (transforms).

What changes

New fields are appended after today's columns, so readers by name or by position keep working:

JSON Lines CSV
javaFrames, javaFrameKinds, kernelFrames, userFrames (arrays, null when absent) — (CSV stays flat)
canonicalJavaStack canonical_java_stack
threadPool thread_pool
run (--run-label, default profile label, else first session id) run
estimateAvailable estimate_available
  • JSON Lines counters are now JSON numbers up to 2^53-1 and strings beyond. --numbers string restores the previous form.
  • --run-metadata FILE writes a single JSON object with the profile's run, sources, sampling, dimensions, estimate validity and totals.
  • CI: the x86-64 glibc job installs DuckDB with brew install duckdb when it isn't already on the runner, and sets JONOFFCPU_REQUIRE_DUCKDB=true so the SQL checks fail instead of skipping there.

Verification

  • ./gradlew spotlessCheck :jonoffcpu-correlator:check -PjonoffcpuFixtures=… passes locally with DuckDB 1.5.5 and JONOFFCPU_REQUIRE_DUCKDB=true.
  • On the broker fixture: 2,791 rows; every javaFrames joined with ; equals javaStack; javaFrameKinds has the same length; observed nanoseconds sum to 8,519,334,220,784; no generated-class address is left in canonicalJavaStack; --run-metadata total matches.
  • DuckDB infers javaFrames as VARCHAR[] and observedNanos as BIGINT with no columns argument.
  • The boundary recipe, adapted to javaFrames, returns internalConsumerFlow 11.982 s / 4,245 and GrowableBatchedArrayBlockingQueue.offer 4.946 s / 1,565. The two-run comparison reproduces the spec's table exactly (Alpine/Wolfi s/M: 2.034/2.396, 0.735/0.989, 0.091/0.161, 0.116/0.124, 0.064/0.101).
  • Compared with the previous export of the same profile, every --numbers string JSONL row and every CSV line starts with the previous row byte for byte.
  • ExportTest: arrays, canonical names, pools (pulsar-io-3-25 → pulsar-io-#-#, [tid=12345] → [tid=#]), the 2^53 boundary, field order, run metadata, and DuckDB type inference plus a recipe.

The README's "Analyzing with SQL" recipes land with the README PR at the top of the stack.

@lhotari
lhotari added this pull request to stack #14 September 23, 2026 20:16
@lhotari
lhotari removed this pull request from stack #14 September 23, 2026 20:17
@lhotari
lhotari added this pull request to stack #15 September 23, 2026 20:22
Base automatically changed from stacks-transforms to main September 23, 2026 20:26
Every table of the Apache Pulsar analysis was computed in DuckDB from export,
and every query started the same way: split the joined stacks twice, cast
the string counters, add a run column by hand and regexp the thread names.
export now carries that itself, appended after the existing columns so
today's readers keep working:

- JSON Lines rows add javaFrames, javaFrameKinds, kernelFrames and
  userFrames as arrays (null for an absent stack); CSV stays flat.
- canonicalJavaStack / canonical_java_stack: the Java stack without
  generated-class addresses, the rule of stacks --canonical-names, so two
  runs' stacks join.
- threadPool / thread_pool: the thread name with digit runs as '#'.
- run: --run-label, by default the profile's label, else its first
  source's session id.
- estimateAvailable / estimate_available on every row, so a query cannot
  use the estimated columns by accident.
- JSON Lines counters are numbers up to 2^53-1 (strings beyond it);
  --numbers string restores the previous form.
- --run-metadata FILE writes the profile's provenance, dimensions,
  estimate validity and totals as one JSON object to join on run.

On the Pulsar broker fixture the export has 2,791 rows whose arrays join
back to javaStack and whose observed nanoseconds sum to 8,519,334,220,784;
DuckDB infers VARCHAR[] and BIGINT without a schema, and the README's
boundary and two-run comparison recipes reproduce the spec's tables
(internalConsumerFlow 11.982 s / 4,245 intervals; 2.034 vs 2.396 s per
million messages on Alpine vs Wolfi). With --numbers string every row
extends the previous export's row byte for byte, as every CSV line does.

ExportTest checks DuckDB's type inference and a recipe when duckdb is on
the path; CI installs it with Homebrew on the x86-64 glibc job and sets
JONOFFCPU_REQUIRE_DUCKDB so the check cannot silently skip there.
@lhotari
lhotari merged commit deb6628 into main Sep 23, 2026
5 checks passed
@lhotari
lhotari deleted the export-for-sql branch September 24, 2026 13:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant