DataFrame 1.0 Help

map

Computes a new value for every value, row, or key–group pair of the receiver, and collects the results into a List, a DataFrame, a DataColumn, or a FrameColumn.

Related operations: Add / map / remove columns

All map operations share the name but differ in what they go over and what they give back:

Operation

Goes over

Returns

map

rows of a DataFrame

List

mapToColumn

rows of a DataFrame

a single DataColumn

mapToFrame

rows of a DataFrame

a new DataFrame

map/mapIndexed

values of a DataColumn

a DataColumn of the same size

map

key–group pairs of a GroupBy

List

mapToRows

key–group pairs of a GroupBy

a DataFrame

mapToFrames

key–group pairs of a GroupBy

a FrameColumn

Each result keeps the order of the values, rows, or key–group pairs it was computed from.

Every example on this page uses the same DataFrame:

df

map

Maps the rows of a DataFrame into a List with one element per row.

map { rowExpression }: List<T> rowExpression: DataRow.(DataRow) -> Value
df.map { 2021 - it.age }

See row expressions

A ColumnGroup is also a DataFrame, so map on a column group is this operation: it goes over the rows of the group and returns a List. To get a DataColumn of the same size instead — a column of the rows of the group — call asDataColumn() on the group first, and then map.

mapToColumn

Maps the rows of a DataFrame into a single new DataColumn with one value per row.

mapToColumn(columnName) { rowExpression }: DataColumn rowExpression: DataRow.(DataRow) -> Value
df.mapToColumn("year of birth") { 2021 - age }
df.mapToColumn("year of birth") { 2021 - "age"<Int>() }

See row expressions

The new column is standalone: the original DataFrame is not changed and does not contain it. Use add to get a DataFrame with the new column in it.

Inside the row expression, prev()?.newValue() gives the value already computed for the preceding row — null for the first row, where there is no preceding one. This is how running totals and other recurrences are expressed; see add.

mapToFrame

Maps the rows of a DataFrame into a new DataFrame made of the described columns.

mapToFrame { columnMapping columnMapping ... } : DataFrame columnMapping = column into columnName | columnName from column | columnName from { rowExpression } | +column
df.mapToFrame { "year of birth" from { 2021 - age } expr { age > 18 } into "is adult" name.lastName.map { it.length } into "last name length" "full name" from { name.firstName + " " + name.lastName } +city }
df.mapToFrame { "year of birth" from { 2021 - "age"<Int>() } expr { "age"<Int>() > 18 } into "is adult" "name"["lastName"]<String>().map { it.length } into "last name length" "full name" from { "name"["firstName"]<String>() + " " + "name"["lastName"]<String>() } +"city" }

The result holds only the described columns, in the order in which they are described. This is what makes mapToFrame different from add, where the columns of the original DataFrame are part of the result as well. In the example above, city is in the result only because of the +city line, and name and age are not in it at all.

map on DataColumn

Maps the values of a DataColumn into a new DataColumn of the same size. mapIndexed gives the position of the value as well, starting at 0.

map { value -> newValue }: DataColumn map(type) { value -> newValue }: DataColumn mapIndexed { index, value -> newValue }: DataColumn mapIndexed(type) { index, value -> newValue }: DataColumn
// A column of last name lengths; it keeps the name of the original column, // so it is renamed here df.name.lastName.map { it.length }.rename("lastNameLength")

The new column has the same name as the original one, so it is usually renamed on the spot or given a name by the operation it is passed to.

mapIndexed also gives the position of the value:

// "1. Alice", "2. Bob", ... df.name.firstName.mapIndexed { i, firstName -> "${i + 1}. $firstName" }

Which kind of column you get follows the type of the new column — the type argument, or the type given explicitly — and not the computed values: a DataFrame type gives a FrameColumn, a DataRow type gives a ColumnGroup, and any other type gives a ValueColumn. A nullable DataFrame type belongs to the last group, because a FrameColumn cannot hold null.

infer only concerns a ValueColumn: it decides whether the type of that column is the given one as it is, or the type of the computed values. For a ColumnGroup and a FrameColumn it changes nothing.

The overloads with an explicit type are for the cases where the type of the new column is only known at runtime. The computed values are put into the column as they are, without any conversion, so the type has to fit them.

A ValueColumn can never have a non-nullable DataFrame type, so a call that would give it one fails with an IllegalArgumentException. That happens when the computed values are dataframes and Infer.Type derives a DataFrame type for them, and also under a nullable DataFrame type when none of the computed values is null: the default Infer.Nulls then drops the nullability and leaves exactly that forbidden type. With at least one null among them the same call succeeds and gives a ValueColumn of the nullable type.

map on GroupBy

Maps the key–group pairs of a GroupBy: every row of a GroupBy is one key–group pair — the key values, and the group of rows that belongs to them (see groupBy).

map { GroupWithKey -> value }: List<Value> mapToRows { GroupWithKey -> DataRow? }: DataFrame mapToFrames { GroupWithKey -> DataFrame }: FrameColumn

The lambda receives the pair as a GroupWithKey, both as the receiver and as the argument, so the key values are available as key (a DataRow) and the rows of the group as group (a DataFrame).

// The number of people per city, as a list, in the order of the groups: [1, 1, 2, 1, 1, 1] df.groupBy { city }.map { group.rowsCount() }
// The oldest person of every city, one row per city df.groupBy { city }.mapToRows { group.sortByDesc { age }.firstOrNull() }
// The two oldest people for each first name, as a frame column: // every frame keeps all the columns of the original, including the first name it was grouped by df.groupBy { name.firstName }.mapToFrames { group.sortByDesc { age }.take(2) }

mapToFrames names the new column after GroupBy.groups ("group" by default); call concat() on it to get all of those dataframes back as one DataFrame. concatWithKeys is built exactly this way.

25 September 2026