Kotlin DataFrame supports reading Apache Parquet files through the Apache Arrow integration.
Requires the dataframe-arrow module, which is included by default in the general dataframe artifact and when using %use dataframe for Kotlin Notebook.
Reading Parquet Files
Kotlin DataFrame provides four readParquet() methods that can read from different source types. All overloads accept optional nullability inference settings and batchSize for Arrow scanning.
// 1) URLs
public fun DataFrame.Companion.readParquet(
vararg urls: URL,
nullability: NullabilityOptions = NullabilityOptions.Infer,
batchSize: Long = ARROW_PARQUET_DEFAULT_BATCH_SIZE,
): AnyFrame
// 2) Strings (interpreted as file paths or URLs, e.g., "data/file.parquet", "file://", or "http(s)://")
public fun DataFrame.Companion.readParquet(
vararg strUrls: String,
nullability: NullabilityOptions = NullabilityOptions.Infer,
batchSize: Long = ARROW_PARQUET_DEFAULT_BATCH_SIZE,
): AnyFrame
// 3) Paths
public fun DataFrame.Companion.readParquet(
vararg paths: Path,
nullability: NullabilityOptions = NullabilityOptions.Infer,
batchSize: Long = ARROW_PARQUET_DEFAULT_BATCH_SIZE,
): AnyFrame
// 4) Files
public fun DataFrame.Companion.readParquet(
vararg files: File,
nullability: NullabilityOptions = NullabilityOptions.Infer,
batchSize: Long = ARROW_PARQUET_DEFAULT_BATCH_SIZE,
): AnyFrame
These overloads are defined in the dataframe-arrow module and internally use FileFormat.PARQUET from Apache Arrow’s Dataset API to scan the data and materialize it as a Kotlin DataFrame.
ARROW_PARQUET_DEFAULT_BATCH_SIZE is 32768 rows — the number of rows Arrow reads per batch while scanning. It is a public constant, so you can reference it when tuning batchSize relative to the default.
Examples
// Read from file paths (as strings)
val df = DataFrame.readParquet("data/sales.parquet")
// Read from Path objects
val df = DataFrame.readParquet(path)
// Read from URLs
val df = DataFrame.readParquet(url)
// Read from File objects
val df = DataFrame.readParquet(file)
The distinction is not cosmetic. With isAdjustedToUTC = true the number counts time units since 1970-01-01T00:00:00Z, so it identifies one point on the time-line — an instant. With isAdjustedToUTC = false it is a bare calendar-and-clock reading with no zone, which identifies no single point in time. This is the case the Parquet specification describes in detail.
Writers such as PyArrow, Polars and pandas set isAdjustedToUTC = true for every time-zone-aware column, so those columns are read as Instant. MILLIS, MICROS and NANOS are all supported, and the full precision is kept — a NANOS column keeps all nine fractional digits.
An Instant names a point on the time-line but no wall clock, so reading one on somebody's clock takes an explicit zone:
Column selection: Because the readParquet method reads all columns, use DataFrame operations like select() immediately after reading to reduce memory usage in later operations
Predicate pushdown: Currently not supported—filtering happens after data is loaded into memory