Skip to content

New query evaluator using DataFusion - #1527

Draft
Tpt wants to merge 5 commits into
mainfrom
fusion
Draft

Tpt wants to merge 5 commits into
mainfrom
fusion

Conversation

@Tpt

@Tpt Tpt commented Nov 30, 2025

Copy link
Copy Markdown
Collaborator

Very WIP, will stay a draft for a long time

@Tpt
Tpt force-pushed the fusion branch 3 times, most recently from 063dc78 to 85e0c51 Compare December 5, 2025 21:15
@codspeed

codspeed Bot commented Dec 11, 2025 •

Copy link
Copy Markdown

Merging this PR will degrade performance by 6.15%

❌ 1 regressed benchmark
✅ 15 untouched benchmarks
⏩ 78 skipped benchmarks1

⚠️ Please fix the performance issues or acknowledge them on CodSpeed.

Performance Changes

Mode Benchmark BASE HEAD Efficiency
❌ Memory RDFC-1.0 (SHA256) 1.1 MB 1.1 MB -6.15%

Comparing fusion (f0d6c84) with main (27ee297)2

Open in CodSpeed

Footnotes

  1. 78 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

  2. No successful run was found on main (654bb80) during the generation of this report, so 27ee297 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

@ashleysommer

Copy link
Copy Markdown

Wow @Tpt this looks amazing. This looks like it was an enormous amount of work.
Do you think this will (eventually) become the default query evaluator?
Are there any downsides, eg, particularly with regard to DataFusion using Apache Arrow in-memory structures, and the whole Oxigraph->Arrow conversion overhead?

@Tpt

Tpt commented Dec 15, 2025

Copy link
Copy Markdown
Collaborator Author

@ashleysommer Thank you! Was much easier than the JSON-LD streaming parser (I have not finished it...). It currently has a few downsides:

  • much larger dependency footprint
  • currently much slower on "BSBM explore 5000 query in memory" it takes 150s vs <1s. My guess is that having better join reordering and use for loop joins where it makes sense might help a lot
  • the conversion to Arrow is indeed very costly. I am considering writing a new storage backend on top of Parquet, likely for read-only workloads first
  • EXISTS support is currently very spotty and it will require significant work to get it done properly

@Tpt

Tpt commented May 17, 2026

Copy link
Copy Markdown
Collaborator Author

Some missing elements in this MR:

  • RocksDB iter is quite slow, making large reads not very good. Tweaking RocksDB evaluation to do less hash/merge joins and more "for loop" joins is needed to keep the same performances on top of RocksDB (and nudges more toward rewriting the storage to be more read-friendly at the cost of maybe being less write-friendly)
  • EXISTS and LATERAL are written as correlated queries. Sadly DataFusion does not know how to evaluate them without decorating the queries first. We should either change the emitted logical plan in something DataFusion is able to decorellate, write a better decorrelation algorithm or add correlated evaluation to DataFusion (this is sorted by personal order of preference).

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants