Skip to content

Avro Codecs improvements #3636

Description

@pomadchin

Avro 1.12.1 caches schemas by default; uses 'fast readers' to keep them in cache. That originally lead to to tests OOM due to too many objects being created and not eventually cleaned up.

i.e. Spark implements own caches, https://github.com/apache/spark/blob/v4.2.0/core/src/main/scala/org/apache/spark/serializer/GenericAvroSerializer.scala#L48-L60

This task is to see if we could make use of Avro fast readers and if we could do better in terms of the Avro write performance. We could at least make the majority of the codecs and their schemas lazy vals

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions