DataFrame 1.0 Help

Apache Arrow

Kotlin DataFrame supports reading from and writing to Apache Arrow files.

Requires the dataframe-arrow module, which is included by default in the general dataframe artifact and in %use dataframe for Kotlin Notebook.

Read

DataFrame supports both the Arrow interprocess streaming format and the Arrow random access format.

You can read a DataFrame from Apache Arrow data sources (via a file path, URL, or stream) using the readArrowFeather() method:

val df = DataFrame.readArrowFeather("example.feather")
val df = DataFrame.readArrowFeather("https://kotlin.github.io/dataframe/resources/example.feather")

Write

A DataFrame can be written to Arrow format using the interprocess streaming or random access format. Output targets include WritableByteChannel, OutputStream, File, or ByteArray.

See Writing to Apache Arrow formats for more details.

Type mapping

Reading

Arrow type

Kotlin type

Null

Nothing?

Bool

Boolean

Int(8, signed)/Int(16, signed)/Int(32, signed)/Int(64, signed)

Byte/Short/Int/Long

Int(8, unsigned)/Int(16, unsigned)/Int(32, unsigned)/Int(64, unsigned)

Short/Int/Long/BigInteger

FloatingPoint(SINGLE)/FloatingPoint(DOUBLE)

Float/Double

Decimal (128- and 256-bit)

BigDecimal

Utf8, LargeUtf8, Utf8View

String

Binary, LargeBinary, BinaryView

ByteArray

Date(DAY)

kotlinx.datetime.LocalDate

Date(MILLISECOND)

kotlinx.datetime.LocalDateTime

Time(SECOND / MILLISECOND / MICROSECOND / NANOSECOND)

kotlinx.datetime.LocalTime

Timestamp(unit, null) — no time zone

kotlinx.datetime.LocalDateTime

Timestamp(unit, tz) — with a time zone

kotlin.time.Instant

Duration

kotlin.time.Duration

Struct

ColumnGroup

List, LargeList

List<T>, or a FrameColumn for a list of structs

Anything else raises NotImplementedError. Column nullability comes from the nullability argument (NullabilityOptions.Infer by default, which marks a column nullable only if it actually contains nulls).

A timestamp with a time zone is an offset from 1970-01-01T00:00:00Z and so identifies a single point on the time-line, which is why it becomes an Instant; a timestamp without one is a bare calendar-and-clock reading that identifies no such point, and stays a LocalDateTime. This is also how Parquet's isAdjustedToUTC flag is mapped — see Timestamps and time zones.

Writing

Kotlin type

Arrow type

Nothing?

Null

String

Utf8

Boolean

Bool

Byte/Short/Int/Long

Int(8 / 16 / 32 / 64, signed)

Float/Double

FloatingPoint(SINGLE)/FloatingPoint(DOUBLE)

LocalDate (kotlinx.datetime or java.time)

Date(DAY)

LocalDateTime (kotlinx.datetime or java.time)

Date(MILLISECOND)

LocalTime (kotlinx.datetime or java.time)

Time(NANOSECOND)

Instant (kotlin.time or java.time)

Timestamp(MICROSECOND, "UTC")

ColumnGroup

Struct

Any other type is written as Utf8 (its toString()), reported through the ConvertingMismatch subscriber. When you supply an explicit target Schema, Timestamp fields are also accepted in every unit, with or without a time zone, and the column is converted accordingly.

An instant outside the target unit's range — the year 2500 in a Timestamp(NANOSECOND, "UTC") field — is reported as ConvertingMismatch.ValueOutOfRange, then refused with a ConvertingException under ArrowWriter.Mode.STRICT (the default) or written as null under ArrowWriter.Mode.LOYAL. Dropping it needs a nullable field, so a non-nullable one is refused in either mode.

Writing a Timestamp field resolves every conversion it needs against UTC, never against the JVM's default time zone — a local date-time to an instant and back, a LocalDate to the start of its day, a number to calendar-and-clock fields (a number is read as epoch milliseconds). The result therefore depends only on the DataFrame and the target schema, not on the environment it is written in.

29 September 2026