Apache Arrow
Kotlin DataFrame supports reading from and writing to Apache Arrow files.
Requires the dataframe-arrow module, which is included by default in the general dataframe artifact and in %use dataframe for Kotlin Notebook.
Read
DataFrame supports both the Arrow interprocess streaming format and the Arrow random access format.
You can read a DataFrame from Apache Arrow data sources (via a file path, URL, or stream) using the readArrowFeather() method:
Write
A DataFrame can be written to Arrow format using the interprocess streaming or random access format. Output targets include WritableByteChannel, OutputStream, File, or ByteArray.
See Writing to Apache Arrow formats for more details.
Type mapping
Reading
Arrow type | Kotlin type |
|---|---|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
| |
| |
| |
| |
|
|
| |
|
|
Anything else raises NotImplementedError. Column nullability comes from the nullability argument (NullabilityOptions.Infer by default, which marks a column nullable only if it actually contains nulls).
A timestamp with a time zone is an offset from 1970-01-01T00:00:00Z and so identifies a single point on the time-line, which is why it becomes an Instant; a timestamp without one is a bare calendar-and-clock reading that identifies no such point, and stays a LocalDateTime. This is also how Parquet's isAdjustedToUTC flag is mapped — see Timestamps and time zones.
Writing
Kotlin type | Arrow type |
|---|---|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
Any other type is written as Utf8 (its toString()), reported through the ConvertingMismatch subscriber. When you supply an explicit target Schema, Timestamp fields are also accepted in every unit, with or without a time zone, and the column is converted accordingly.
An instant outside the target unit's range — the year 2500 in a Timestamp(NANOSECOND, "UTC") field — is reported as ConvertingMismatch.ValueOutOfRange, then refused with a ConvertingException under ArrowWriter.Mode.STRICT (the default) or written as null under ArrowWriter.Mode.LOYAL. Dropping it needs a nullable field, so a non-nullable one is refused in either mode.
Writing a Timestamp field resolves every conversion it needs against UTC, never against the JVM's default time zone — a local date-time to an instant and back, a LocalDate to the start of its day, a number to calendar-and-clock fields (a number is read as epoch milliseconds). The result therefore depends only on the DataFrame and the target schema, not on the environment it is written in.