Repository navigation
Make export friendlier to SQL - #11
Merged
Merged
Conversation
This was referenced Sep 23, 2026
lhotari
added this pull request to stack #14
September 23, 2026 20:16
lhotari
removed this pull request from stack #14
September 23, 2026 20:17
lhotari
force-pushed
the
export-for-sql
branch
from
September 23, 2026 20:22
7a4a830 to
9861a9f
Compare
lhotari
added this pull request to stack #15
September 23, 2026 20:22
Every table of the Apache Pulsar analysis was computed in DuckDB from export, and every query started the same way: split the joined stacks twice, cast the string counters, add a run column by hand and regexp the thread names. export now carries that itself, appended after the existing columns so today's readers keep working: - JSON Lines rows add javaFrames, javaFrameKinds, kernelFrames and userFrames as arrays (null for an absent stack); CSV stays flat. - canonicalJavaStack / canonical_java_stack: the Java stack without generated-class addresses, the rule of stacks --canonical-names, so two runs' stacks join. - threadPool / thread_pool: the thread name with digit runs as '#'. - run: --run-label, by default the profile's label, else its first source's session id. - estimateAvailable / estimate_available on every row, so a query cannot use the estimated columns by accident. - JSON Lines counters are numbers up to 2^53-1 (strings beyond it); --numbers string restores the previous form. - --run-metadata FILE writes the profile's provenance, dimensions, estimate validity and totals as one JSON object to join on run. On the Pulsar broker fixture the export has 2,791 rows whose arrays join back to javaStack and whose observed nanoseconds sum to 8,519,334,220,784; DuckDB infers VARCHAR[] and BIGINT without a schema, and the README's boundary and two-run comparison recipes reproduce the spec's tables (internalConsumerFlow 11.982 s / 4,245 intervals; 2.034 vs 2.396 s per million messages on Alpine vs Wolfi). With --numbers string every row extends the previous export's row byte for byte, as every CSV line does. ExportTest checks DuckDB's type inference and a recipe when duckdb is on the path; CI installs it with Homebrew on the x86-64 glibc job and sets JONOFFCPU_REQUIRE_DUCKDB so the check cannot silently skip there.
lhotari
force-pushed
the
export-for-sql
branch
from
September 23, 2026 20:26
9861a9f to
4e6f40c
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Second PR of the 0.5.0 correlator stack (GitHub stack #15; spec:
export-for-sql.md). Based on #10 (transforms).What changes
New fields are appended after today's columns, so readers by name or by position keep working:
javaFrames,javaFrameKinds,kernelFrames,userFrames(arrays,nullwhen absent)canonicalJavaStackcanonical_java_stackthreadPoolthread_poolrun(--run-label, default profile label, else first session id)runestimateAvailableestimate_available--numbers stringrestores the previous form.--run-metadata FILEwrites a single JSON object with the profile's run, sources, sampling, dimensions, estimate validity and totals.brew install duckdbwhen it isn't already on the runner, and setsJONOFFCPU_REQUIRE_DUCKDB=trueso the SQL checks fail instead of skipping there.Verification
./gradlew spotlessCheck :jonoffcpu-correlator:check -PjonoffcpuFixtures=…passes locally with DuckDB 1.5.5 andJONOFFCPU_REQUIRE_DUCKDB=true.javaFramesjoined with;equalsjavaStack;javaFrameKindshas the same length; observed nanoseconds sum to 8,519,334,220,784; no generated-class address is left incanonicalJavaStack;--run-metadatatotal matches.javaFramesasVARCHAR[]andobservedNanosasBIGINTwith nocolumnsargument.javaFrames, returnsinternalConsumerFlow11.982 s / 4,245 andGrowableBatchedArrayBlockingQueue.offer4.946 s / 1,565. The two-run comparison reproduces the spec's table exactly (Alpine/Wolfi s/M: 2.034/2.396, 0.735/0.989, 0.091/0.161, 0.116/0.124, 0.064/0.101).--numbers stringJSONL row and every CSV line starts with the previous row byte for byte.ExportTest: arrays, canonical names, pools (pulsar-io-3-25→pulsar-io-#-#,[tid=12345]→[tid=#]), the 2^53 boundary, field order, run metadata, and DuckDB type inference plus a recipe.The README's "Analyzing with SQL" recipes land with the README PR at the top of the stack.