Skip to content

Add batch/columnar write API for Arrow VectorSchemaRoot (Java parity with C++/Python) #3733

Description

@qzyu999

Problem

parquet-java has no public API to write Arrow VectorSchemaRoot to Parquet files. Callers must construct row objects and feed them through ParquetWriter<T>.write(T) one at a time. Arrow C++ and PyArrow support this natively via parquet::WriteTable() / pq.write_table().

Related prior discussion:

Motivation

Several downstream Java projects work with Arrow-columnar data internally and must materialize row objects solely to satisfy ParquetWriter's input API:

  • Apache Iceberg (FileAppender<Record>) — iceberg#17748
  • Apache Fluss (Arrow-native streaming storage) — fluss#4047
  • Apache Paimon (worked around this by building their own paimon-arrow writer that bypasses parquet-java's row API entirely)

Key Insight

For PLAIN-encoded, fixed-width columns, Arrow's in-memory format (contiguous little-endian values) is identical to Parquet's PLAIN page encoding. A zero-copy path is possible by wrapping the Arrow data buffer directly as a BytesInput and passing it to PageWriter.writePage().

For nullable columns, Arrow's validity bitmap can be scanned for contiguous non-null runs, with each run bulk-copied and definition levels encoded as RLE runs of the same value — O(null_transitions) instead of O(N).

Proposed Design

A new ArrowParquetWriter in the parquet-arrow module that writes pages directly to PageWriter rather than going through RecordConsumer/ColumnWriter:

ArrowParquetWriter.writeBatch(VectorSchemaRoot)
    → per column: selects optimal write strategy
    → writes pages to PageWriter via getPageWriter(ColumnDescriptor)
    → manages row groups via ParquetFileWriter

Per-column strategy selection (best to worst):

  1. Zero-copy (non-null + fixed-width + PLAIN): wrap Arrow buffer as BytesInput directly. RL/DL as single-value RLE runs. Stats from sequential buffer scan.
  2. Bulk-copy with nulls (nullable + fixed-width + PLAIN): scan validity bitmap for non-null runs, bulk-copy each, emit DL as RLE runs per null-transition.
  3. Variable-width rewrite (PLAIN + string/binary): single-pass transformation of Arrow offset+data buffers to Parquet's length-prefixed format.
  4. Dictionary: map Arrow's DictionaryEncodedVector to Parquet dictionary pages, or build dictionary from plain vector.
  5. Fallback: per-value through ValuesWriter (same cost as today, for unsupported encoding/type combos).

All verified public APIs exist for this:

  • PageWriteStore.getPageWriter(ColumnDescriptor) — get column page writer directly
  • PageWriter.writePage(BytesInput, valueCount, rowCount, stats, encodings) — write pre-encoded pages
  • BytesInput.from(ByteBuffer) — zero-copy buffer wrapping
  • RunLengthBitPackingHybridEncoder — encode RL/DL levels
  • ColumnChunkPageWriteStore.flushToFileWriter() — flush pages to file

Why Not Extend ParquetWriter?

ParquetWriter.write(T) calls InternalParquetRecordWriter.write() which increments recordCount by 1 per call and checks row-group boundaries based on that count. This is incompatible with batch semantics. The writer must manage ParquetFileWriter and row groups directly.

Scope

  • Parquet file format unchanged
  • Existing ParquetWriter<T> API unchanged
  • parquet-arrow gains parquet-hadoop as a compile dependency (for ParquetFileWriter, ColumnChunkPageWriteStore)
  • No Hadoop runtime dependency for uncompressed output (compression requires caller-supplied CompressionCodecFactory)
  • Flat schemas (primitive columns) first; nested types deferred

Implementation Phases

  1. Core infrastructure + zero-copy for non-null fixed-width PLAIN columns
  2. Nullable column support (validity bitmap scanning + run-based bulk copy)
  3. Variable-width types (string/binary offset rewriting)
  4. Dictionary encoding support
  5. Bloom filter, nested types, advanced features

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions