dataiter.DataFrame

__init__() aggregate() anti_join() cbind() colnames columns compare() copy() count() deepcopy() drop_na() filter() filter_out() from_arrow() from_json() from_pandas() full_join() group_by() head() inner_join() left_join() map() modify() ncol nrow pivot_longer() pivot_wider() print_() print_memory_use() print_na_counts() rbind() read_csv() read_json() read_npz() read_parquet() read_pickle() rename() sample() select() semi_join() slice() slice_off() sort() split() tail() to_arrow() to_json() to_list_of_dicts() to_pandas() to_string() unique() unselect() update() write_csv() write_json() write_npz() write_parquet() write_pickle()

class dataiter.DataFrame(*args, **kwargs)[source]

A class for tabular data.

DataFrame is a subclass of dict, with columns being DataFrameColumn, which are Vector, which are NumPy ndarray. This means that basic dict methods, such as items(), keys() and values() can be used iterate over and manage the data as a whole and NumPy functions and array methods can be used for fast vectorized computations on the data.

Columns can be accessed by attribute notation, e.g. data.x in addition to data["x"]. In most cases, attribute access should be more convenient and is the way recommended by dataiter. You’ll still need to use the bracket notation for any column names that are not valid identifiers, such as ones with spaces, or ones that conflict with dict methods, such as “items”.

DataFrame does not support indexing directly as the bracket notation is used to refer to dict keys, i.e. columns by name. If you want to index the whole data frame object, use the method slice(). Individual columns are indexed the same as NumPy arrays.

__init__(*args, **kwargs)[source]

Return a new data frame.

args and kwargs are like for dict.

https://docs.python.org/3/library/stdtypes.html#dict

aggregate(**colname_function_pairs)[source]

Return group-wise calculated summaries.

Usually aggregation is preceded by grouping, which can be conveniently written via method chaining as data.group_by(...).aggregate(...).

In colname_function_pairs, function receives as an argument a data frame object, a group-wise subset of all rows. It should return a scalar value. Common aggregation functions have shorthand helpers available under dataiter, see the guide on aggregation for details.

>>> data = di.read_csv("data/listings.csv")
>>> # The below aggregations are identical. Usually you'll get by
>>> # with the shorthand helpers, but for complicated calculations,
>>> # you might need custom lambda functions.
>>> data.group_by("hood").aggregate(n=di.count(), price=di.mean("price"))
.
           hood     n   price
         string int64 float64
  ───────────── ───── ───────
0         Bronx  1198  90.176
1      Brooklyn 19931 125.056
2     Manhattan 21963 218.855
3        Queens  6068  99.745
4 Staten Island   370 116.908
.
>>> data.group_by("hood").aggregate(n=lambda x: x.nrow, price=lambda x: x.price.mean())
.
           hood     n   price
         string int64 float64
  ───────────── ───── ───────
0         Bronx  1198  90.176
1      Brooklyn 19931 125.056
2     Manhattan 21963 218.855
3        Queens  6068  99.745
4 Staten Island   370 116.908
.
anti_join(other, *by)[source]

Return rows with no matches in other.

by are column names, by which to look for matching rows, or tuples of column names if the correspoding column name differs between self and other.

>>> # All listings that don't have reviews
>>> listings = di.read_csv("data/listings.csv")
>>> reviews = di.read_csv("data/listings-reviews.csv")
>>> listings.anti_join(reviews, "id")
.
     id      hood zipcode guests    sqft price
  int64    string  string  int64 float64 int64
  ───── ───────── ─────── ────── ─────── ─────
0  2060 Manhattan   10040      2     nan   100
1  5136  Brooklyn   11232      4     nan   253
2  7750 Manhattan   10029      1     750    35
3  7801  Brooklyn   11211      4     nan   299
4  8700 Manhattan   10034      2     700    80
5 11943  Brooklyn   11226      1     nan   150
6 15396 Manhattan   10001      4     nan   400
7 16458  Brooklyn   11215      4     nan   225
8 20300 Manhattan   10009      2     nan    50
9 21644 Manhattan   10031      1     nan    89
.
... 30511 rows total
cbind(*others)[source]

Return data frame with columns from others added.

>>> data = di.read_csv("data/listings.csv")
>>> data.cbind(di.DataFrame(x=1))
.
     id      hood zipcode guests    sqft price     x
  int64    string  string  int64 float64 int64 int64
  ───── ───────── ─────── ────── ─────── ───── ─────
0  2060 Manhattan   10040      2     nan   100     1
1  2595 Manhattan   10018      2     nan   225     1
2  3831  Brooklyn   11238      3     500    89     1
3  5099 Manhattan   10016      2     nan   200     1
4  5121  Brooklyn   11216      2     nan    60     1
5  5136  Brooklyn   11232      4     nan   253     1
6  5178 Manhattan   10019      2     nan    79     1
7  5203 Manhattan   10025      1     nan    79     1
8  5238 Manhattan   10002      2     nan   150     1
9  5441 Manhattan   10036      2     nan    99     1
.
... 49530 rows total
property colnames

Get or set column names as a list.

>>> data = di.read_csv("data/listings.csv")
>>> data.head()
.
     id      hood zipcode guests    sqft price
  int64    string  string  int64 float64 int64
  ───── ───────── ─────── ────── ─────── ─────
0  2060 Manhattan   10040      2     nan   100
1  2595 Manhattan   10018      2     nan   225
2  3831  Brooklyn   11238      3     500    89
3  5099 Manhattan   10016      2     nan   200
4  5121  Brooklyn   11216      2     nan    60
5  5136  Brooklyn   11232      4     nan   253
6  5178 Manhattan   10019      2     nan    79
7  5203 Manhattan   10025      1     nan    79
8  5238 Manhattan   10002      2     nan   150
9  5441 Manhattan   10036      2     nan    99
.
>>> data.colnames
['id', 'hood', 'zipcode', 'guests', 'sqft', 'price']
>>> data.colnames = ["a", "b", "c", "d", "e", "f"]
>>> data.head()
.
      a         b      c     d       e     f
  int64    string string int64 float64 int64
  ───── ───────── ────── ───── ─────── ─────
0  2060 Manhattan  10040     2     nan   100
1  2595 Manhattan  10018     2     nan   225
2  3831  Brooklyn  11238     3     500    89
3  5099 Manhattan  10016     2     nan   200
4  5121  Brooklyn  11216     2     nan    60
5  5136  Brooklyn  11232     4     nan   253
6  5178 Manhattan  10019     2     nan    79
7  5203 Manhattan  10025     1     nan    79
8  5238 Manhattan  10002     2     nan   150
9  5441 Manhattan  10036     2     nan    99
.
property columns

Return columns as a list.

compare(other, *by, ignore_columns=[], max_changed=inf)[source]

Find differences against another data frame.

by are identifier columns which are used to uniquely identify rows and match them between self and other. compare will not work if your data lacks suitable identifiers. ignore_columns is an optional list of columns, differences in which to ignore.

compare returns three data frames: added rows, removed rows and changed values. The first two are basically subsets of the rows of self and other, respectively. Changed values are returned as a data frame with one row per differing value (not per differing row). Listing changes will terminate once max_changed is reached.

Warning

compare is experimental, do not rely on it reporting all of the differences correctly. Do not try to give it two huge data frames with very little in common, unless also giving some sensible value for max_changed.

>>> old = di.read_csv("data/vehicles.csv")
>>> new = old.modify(hwy=lambda x: np.minimum(100, x.hwy))
>>> added, removed, changed = new.compare(old, "id")
>>> changed
.
     id column xvalue yvalue
  int64 string  int64  int64
  ───── ────── ────── ──────
0 33640    hwy    100    109
1 33396    hwy    100    108
2 34392    hwy    100    108
3 33265    hwy    100    105
4 33905    hwy    100    105
5 33558    hwy    100    102
6 34699    hwy    100    101
7 34918    hwy    100    101
8 33307    hwy    100    105
.
copy()[source]

Return a shallow copy.

count(*colnames)[source]

Return row counts grouped by colnames.

>>> data = di.read_csv("data/listings.csv")
>>> data.count("hood")
.
           hood     n
         string int64
  ───────────── ─────
0         Bronx  1198
1      Brooklyn 19931
2     Manhattan 21963
3        Queens  6068
4 Staten Island   370
.
deepcopy()[source]

Return a deep copy.

drop_na(*colnames)[source]

Return data frame without rows that have missing values in colnames.

>>> data = di.read_csv("data/listings.csv")
>>> data.drop_na("sqft")
.
     id      hood zipcode guests    sqft price
  int64    string  string  int64 float64 int64
  ───── ───────── ─────── ────── ─────── ─────
0  3831  Brooklyn   11238      3     500    89
1  6848  Brooklyn   11211      3     500   140
2  7750 Manhattan   10029      1     750    35
3  8490  Brooklyn   11216      5     800   120
4  8700 Manhattan   10034      2     700    80
5  9704 Manhattan   10027      2     900    52
6 12343 Manhattan   10027      3       1   150
7 13050  Brooklyn   11221      5    1400   120
8 16974 Manhattan   10035      8    2200   225
9 17747  Brooklyn   11238      2    1000   105
.
... 396 rows total
filter(rows=None, **colname_value_pairs)[source]

Return rows that match condition.

Filtering can be done by either rows or colname_value_pairs. rows can be either a boolean vector or a function that receives the data frame as argument and returns a boolean vector. The latter is especially useful in a method chaining context where you don’t have direct access to the data frame in question. Alternatively, colname_value_pairs provides a shorthand to check against a fixed value. See the example below of equivalent filtering all three ways.

>>> data = di.read_csv("data/listings.csv")
>>> data.filter((data.hood == "Manhattan") & (data.guests == 2))
.
     id      hood zipcode guests    sqft price
  int64    string  string  int64 float64 int64
  ───── ───────── ─────── ────── ─────── ─────
0  2060 Manhattan   10040      2     nan   100
1  2595 Manhattan   10018      2     nan   225
2  5099 Manhattan   10016      2     nan   200
3  5178 Manhattan   10019      2     nan    79
4  5238 Manhattan   10002      2     nan   150
5  5441 Manhattan   10036      2     nan    99
6  5552 Manhattan   10014      2     nan   160
7  8700 Manhattan   10034      2     700    80
8  9668 Manhattan   10031      2     nan    50
9  9704 Manhattan   10027      2     900    52
.
... 10209 rows total
>>> data.filter(lambda x: (x.hood == "Manhattan") & (x.guests == 2))
.
     id      hood zipcode guests    sqft price
  int64    string  string  int64 float64 int64
  ───── ───────── ─────── ────── ─────── ─────
0  2060 Manhattan   10040      2     nan   100
1  2595 Manhattan   10018      2     nan   225
2  5099 Manhattan   10016      2     nan   200
3  5178 Manhattan   10019      2     nan    79
4  5238 Manhattan   10002      2     nan   150
5  5441 Manhattan   10036      2     nan    99
6  5552 Manhattan   10014      2     nan   160
7  8700 Manhattan   10034      2     700    80
8  9668 Manhattan   10031      2     nan    50
9  9704 Manhattan   10027      2     900    52
.
... 10209 rows total
>>> data.filter(hood="Manhattan", guests=2)
.
     id      hood zipcode guests    sqft price
  int64    string  string  int64 float64 int64
  ───── ───────── ─────── ────── ─────── ─────
0  2060 Manhattan   10040      2     nan   100
1  2595 Manhattan   10018      2     nan   225
2  5099 Manhattan   10016      2     nan   200
3  5178 Manhattan   10019      2     nan    79
4  5238 Manhattan   10002      2     nan   150
5  5441 Manhattan   10036      2     nan    99
6  5552 Manhattan   10014      2     nan   160
7  8700 Manhattan   10034      2     700    80
8  9668 Manhattan   10031      2     nan    50
9  9704 Manhattan   10027      2     900    52
.
... 10209 rows total
filter_out(rows=None, **colname_value_pairs)[source]

Return rows that don’t match condition.

Filtering can be done by either rows or colname_value_pairs. rows can be either a boolean vector or a function that receives the data frame as argument and returns a boolean vector. The latter is especially useful in a method chaining context where you don’t have direct access to the data frame in question. Alternatively, colname_value_pairs provides a shorthand to check against a fixed value. See the example below of equivalent filtering all three ways.

>>> data = di.read_csv("data/listings.csv")
>>> data.filter_out(data.hood == "Manhattan")
.
     id     hood zipcode guests    sqft price
  int64   string  string  int64 float64 int64
  ───── ──────── ─────── ────── ─────── ─────
0  3831 Brooklyn   11238      3     500    89
1  5121 Brooklyn   11216      2     nan    60
2  5136 Brooklyn   11232      4     nan   253
3  5803 Brooklyn   11215      2     nan    89
4  6848 Brooklyn   11211      3     500   140
5  7097 Brooklyn   11205      4     nan   199
6  7801 Brooklyn   11211      4     nan   299
7  8490 Brooklyn   11216      5     800   120
8 10452 Brooklyn   11238      3     nan    70
9 10962 Brooklyn   11215      2     nan    89
.
... 27567 rows total
>>> data.filter_out(lambda x: x.hood == "Manhattan")
.
     id     hood zipcode guests    sqft price
  int64   string  string  int64 float64 int64
  ───── ──────── ─────── ────── ─────── ─────
0  3831 Brooklyn   11238      3     500    89
1  5121 Brooklyn   11216      2     nan    60
2  5136 Brooklyn   11232      4     nan   253
3  5803 Brooklyn   11215      2     nan    89
4  6848 Brooklyn   11211      3     500   140
5  7097 Brooklyn   11205      4     nan   199
6  7801 Brooklyn   11211      4     nan   299
7  8490 Brooklyn   11216      5     800   120
8 10452 Brooklyn   11238      3     nan    70
9 10962 Brooklyn   11215      2     nan    89
.
... 27567 rows total
>>> data.filter_out(hood="Manhattan")
.
     id     hood zipcode guests    sqft price
  int64   string  string  int64 float64 int64
  ───── ──────── ─────── ────── ─────── ─────
0  3831 Brooklyn   11238      3     500    89
1  5121 Brooklyn   11216      2     nan    60
2  5136 Brooklyn   11232      4     nan   253
3  5803 Brooklyn   11215      2     nan    89
4  6848 Brooklyn   11211      3     500   140
5  7097 Brooklyn   11205      4     nan   199
6  7801 Brooklyn   11211      4     nan   299
7  8490 Brooklyn   11216      5     800   120
8 10452 Brooklyn   11238      3     nan    70
9 10962 Brooklyn   11215      2     nan    89
.
... 27567 rows total
classmethod from_arrow(data, *, dtypes={})[source]

Return a new data frame from pyarrow.Table data.

dtypes is an optional dict mapping column names to NumPy datatypes.

classmethod from_json(string, *, columns=[], dtypes={}, **kwargs)[source]

Return a new data frame from JSON string.

columns is an optional list of columns to limit to. dtypes is an optional dict mapping column names to NumPy datatypes. kwargs are passed to json.load.

classmethod from_pandas(data, *, dtypes={})[source]

Return a new data frame from pandas.DataFrame data.

dtypes is an optional dict mapping column names to NumPy datatypes.

full_join(other, *by)[source]

Return data frame with matching rows merged from self and other.

full_join keeps all rows from both data frames, merging matching ones. If there are multiple matches, the first one will be used. For rows, for which matches are not found, missing values are added.

by are column names, by which to look for matching rows, or tuples of column names if the correspoding column name differs between self and other.

>>> listings = di.read_csv("data/listings.csv")
>>> reviews = di.read_csv("data/listings-reviews.csv")
>>> listings.full_join(reviews, "id")
.
     id      hood zipcode guests    sqft price reviews  rating
  int64    string  string  int64 float64 int64 float64 float64
  ───── ───────── ─────── ────── ─────── ───── ─────── ───────
0  2060 Manhattan   10040      2     nan   100     nan     nan
1  2595 Manhattan   10018      2     nan   225      48      94
2  3831  Brooklyn   11238      3     500    89     322      89
3  5099 Manhattan   10016      2     nan   200      78      90
4  5121  Brooklyn   11216      2     nan    60      50      90
5  5136  Brooklyn   11232      4     nan   253     nan     nan
6  5178 Manhattan   10019      2     nan    79     473      84
7  5203 Manhattan   10025      1     nan    79     118      98
8  5238 Manhattan   10002      2     nan   150     161      94
9  5441 Manhattan   10036      2     nan    99     213      97
.
... 49530 rows total
group_by(*colnames)[source]

Return data frame with colnames set for grouped operations, such as aggregate().

head(n=None)[source]

Return the first n rows.

>>> data = di.read_csv("data/listings.csv")
>>> data.head(5)
.
     id      hood zipcode guests    sqft price
  int64    string  string  int64 float64 int64
  ───── ───────── ─────── ────── ─────── ─────
0  2060 Manhattan   10040      2     nan   100
1  2595 Manhattan   10018      2     nan   225
2  3831  Brooklyn   11238      3     500    89
3  5099 Manhattan   10016      2     nan   200
4  5121  Brooklyn   11216      2     nan    60
.
inner_join(other, *by)[source]

Return data frame with matching rows merged from self and other.

inner_join keeps only rows found in both data frames, merging matching ones. If there are multiple matches, the first one will be used.

by are column names, by which to look for matching rows, or tuples of column names if the correspoding column name differs between self and other.

>>> listings = di.read_csv("data/listings.csv")
>>> reviews = di.read_csv("data/listings-reviews.csv")
>>> listings.inner_join(reviews, "id")
.
     id      hood zipcode guests    sqft price reviews  rating
  int64    string  string  int64 float64 int64   int64 float64
  ───── ───────── ─────── ────── ─────── ───── ─────── ───────
0  2595 Manhattan   10018      2     nan   225      48      94
1  3831  Brooklyn   11238      3     500    89     322      89
2  5099 Manhattan   10016      2     nan   200      78      90
3  5121  Brooklyn   11216      2     nan    60      50      90
4  5178 Manhattan   10019      2     nan    79     473      84
5  5203 Manhattan   10025      1     nan    79     118      98
6  5238 Manhattan   10002      2     nan   150     161      94
7  5441 Manhattan   10036      2     nan    99     213      97
8  5552 Manhattan   10014      2     nan   160      66      97
9  5803  Brooklyn   11215      2     nan    89     180      94
.
... 19019 rows total
left_join(other, *by)[source]

Return data frame with matching rows merged from self and other.

left_join keeps all rows in self, merging matching ones. If there are multiple matches, the first one will be used. For rows, for which matches are not found, missing values are added.

by are column names, by which to look for matching rows, or tuples of column names if the correspoding column name differs between self and other.

>>> listings = di.read_csv("data/listings.csv")
>>> reviews = di.read_csv("data/listings-reviews.csv")
>>> listings.left_join(reviews, "id")
.
     id      hood zipcode guests    sqft price reviews  rating
  int64    string  string  int64 float64 int64 float64 float64
  ───── ───────── ─────── ────── ─────── ───── ─────── ───────
0  2060 Manhattan   10040      2     nan   100     nan     nan
1  2595 Manhattan   10018      2     nan   225      48      94
2  3831  Brooklyn   11238      3     500    89     322      89
3  5099 Manhattan   10016      2     nan   200      78      90
4  5121  Brooklyn   11216      2     nan    60      50      90
5  5136  Brooklyn   11232      4     nan   253     nan     nan
6  5178 Manhattan   10019      2     nan    79     473      84
7  5203 Manhattan   10025      1     nan    79     118      98
8  5238 Manhattan   10002      2     nan   150     161      94
9  5441 Manhattan   10036      2     nan    99     213      97
.
... 49530 rows total
map(function)[source]

Apply function to each row in data.

function receives as arguments the full data frame and the loop index. The return value will be a list of whatever function returns.

Note that map is an inefficient method as it iterates over rows instead of doing vectorized computation. map is mostly intended for complicated conditional cases that are difficult to express in vectorized form.

>>> data = di.read_csv("data/listings-reviews.csv")
>>> data.map(lambda x, i: (x.reviews[i], x.rating[i]))[:3]
[(np.int64(48), np.float64(94.0)), (np.int64(322), np.float64(89.0)), (np.int64(78), np.float64(90.0))]
modify(**colname_value_pairs)[source]

Return data frame with columns modified.

In colname_value_pairs, value can be either a vector or a function that receives the data frame as argument and returns a vector. See the example below of equivalent modification with both ways.

Note that column modification can often be done simpler with a plain assignment, such as data.price_per_guest = data.price / data.guests. modify just allows you to do the same in a method chain context.

>>> data = di.read_csv("data/listings.csv")
>>> data.modify(price_per_guest=data.price/data.guests)
.
     id      hood zipcode guests    sqft price price_per_guest
  int64    string  string  int64 float64 int64         float64
  ───── ───────── ─────── ────── ─────── ───── ───────────────
0  2060 Manhattan   10040      2     nan   100          50.000
1  2595 Manhattan   10018      2     nan   225         112.500
2  3831  Brooklyn   11238      3     500    89          29.667
3  5099 Manhattan   10016      2     nan   200         100.000
4  5121  Brooklyn   11216      2     nan    60          30.000
5  5136  Brooklyn   11232      4     nan   253          63.250
6  5178 Manhattan   10019      2     nan    79          39.500
7  5203 Manhattan   10025      1     nan    79          79.000
8  5238 Manhattan   10002      2     nan   150          75.000
9  5441 Manhattan   10036      2     nan    99          49.500
.
... 49530 rows total
>>> data.modify(price_per_guest=lambda x: x.price / x.guests)
.
     id      hood zipcode guests    sqft price price_per_guest
  int64    string  string  int64 float64 int64         float64
  ───── ───────── ─────── ────── ─────── ───── ───────────────
0  2060 Manhattan   10040      2     nan   100          50.000
1  2595 Manhattan   10018      2     nan   225         112.500
2  3831  Brooklyn   11238      3     500    89          29.667
3  5099 Manhattan   10016      2     nan   200         100.000
4  5121  Brooklyn   11216      2     nan    60          30.000
5  5136  Brooklyn   11232      4     nan   253          63.250
6  5178 Manhattan   10019      2     nan    79          39.500
7  5203 Manhattan   10025      1     nan    79          79.000
8  5238 Manhattan   10002      2     nan   150          75.000
9  5441 Manhattan   10036      2     nan    99          49.500
.
... 49530 rows total

If the data frame is grouped, then colname_value_pairs need to be functions, which are applied to group-wise subsets of the data frame. A common use for this is calculating group-wise fractions.

>>> data = di.DataFrame(g=[1, 2, 2, 3, 3, 3])
>>> data.group_by("g").modify(f=lambda x: 1 / x.nrow)
.
      g       f
  int64 float64
  ───── ───────
0     1 1.00000
1     2 0.50000
2     2 0.50000
3     3 0.33333
4     3 0.33333
5     3 0.33333
.
property ncol

Return the amount of columns.

>>> data = di.read_csv("data/listings.csv")
>>> data.ncol
6
property nrow

Return the amount of rows.

>>> data = di.read_csv("data/listings.csv")
>>> data.nrow
49530
pivot_longer(*, ids=None, names=None, values=None)[source]

Pivot data from wide to long format.

ids should be the name or a list of names of identifier columns. All other columns are considered variable columns and will be pivoted. names and values are names of columns into which the variable names and values are put in the result.

>>> wide = di.read_csv("data/listings.csv")
>>> wide
.
     id      hood zipcode guests    sqft price
  int64    string  string  int64 float64 int64
  ───── ───────── ─────── ────── ─────── ─────
0  2060 Manhattan   10040      2     nan   100
1  2595 Manhattan   10018      2     nan   225
2  3831  Brooklyn   11238      3     500    89
3  5099 Manhattan   10016      2     nan   200
4  5121  Brooklyn   11216      2     nan    60
5  5136  Brooklyn   11232      4     nan   253
6  5178 Manhattan   10019      2     nan    79
7  5203 Manhattan   10025      1     nan    79
8  5238 Manhattan   10002      2     nan   150
9  5441 Manhattan   10036      2     nan    99
.
... 49530 rows total
>>> wide.pivot_longer(ids="id", names="name", values="value")
.
       id    name     value
  float64  string    object
  ─────── ─────── ─────────
0    2060    hood Manhattan
1    2060 zipcode     10040
2    2060  guests         2
3    2060    sqft      None
4    2060   price       100
5    2595    hood Manhattan
6    2595 zipcode     10018
7    2595  guests         2
8    2595    sqft      None
9    2595   price       225
.
... 247650 rows total
pivot_wider(*, ids=None, names=None, values=None, rename=None)[source]

Pivot data from long to wide format.

ids should be the name or a list of names of identifier columns by which the result will be unique. If ids is not given, it defaults to all columns except names and values. names should be the name of the column that contains the names of variables, which become new columns. values should be the name of the column that contains the corresponding values. rename is an optional function that you can use to e.g. lowercase the new column names or add a prefix or suffix.

>>> long = di.read_csv("data/downloads.csv")
>>> long = long.sort(date=1)
>>> long
.
  category          date downloads
    string datetime64[D]     int64
  ──────── ───────────── ─────────
0   Darwin    2019-09-16     35106
1    Linux    2019-09-16   2977751
2     null    2019-09-16     70379
3    other    2019-09-16       359
4  Windows    2019-09-16     62458
5   Darwin    2019-09-17     32484
6    Linux    2019-09-17   3174634
7     null    2019-09-17     77128
8    other    2019-09-17       456
9  Windows    2019-09-17     67484
.
... 905 rows total
>>> long.pivot_wider(ids="date", names="category", values="downloads", rename=str.lower)
.
           date  darwin   linux    null   other windows
  datetime64[D] float64 float64 float64 float64 float64
  ───────────── ─────── ─────── ─────── ─────── ───────
0    2019-09-16   35106 2977751   70379     359   62458
1    2019-09-17   32484 3174634   77128     456   67484
2    2019-09-18   34925 3223932   77630     503   67380
3    2019-09-19   40803 3211464   77706     398   73433
4    2019-09-20   41015 3172673   71622     371   80704
5    2019-09-21   28325 2265181   40671     110   38924
6    2019-09-22   27781 2169643   41261     158   34280
7    2019-09-23   40712 3111145   72444     289   72159
8    2019-09-24   57127 3266033   78446     354   82215
9    2019-09-25   87887 3237271   76867     363   82216
.
... 181 rows total
print_(*, max_rows=None, max_width=None, truncate_width=None)[source]

Print data frame to sys.stdout.

print_ does the same as calling Python’s builtin print function, but since it’s a method, you can use it at the end of a method chain instead of wrapping a print call around the whole chain.

>>> di.read_csv("data/listings.csv").print_()
.
     id      hood zipcode guests    sqft price
  int64    string  string  int64 float64 int64
  ───── ───────── ─────── ────── ─────── ─────
0  2060 Manhattan   10040      2     nan   100
1  2595 Manhattan   10018      2     nan   225
2  3831  Brooklyn   11238      3     500    89
3  5099 Manhattan   10016      2     nan   200
4  5121  Brooklyn   11216      2     nan    60
5  5136  Brooklyn   11232      4     nan   253
6  5178 Manhattan   10019      2     nan    79
7  5203 Manhattan   10025      1     nan    79
8  5238 Manhattan   10002      2     nan   150
9  5441 Manhattan   10036      2     nan    99
.
... 49530 rows total
None
print_memory_use()[source]

Print memory use by column and total.

>>> data = di.read_csv("data/listings.csv")
>>> data.print_memory_use()
.
   COLUMN                     DTYPE ITEM_SIZE TOTAL_SIZE
   string                    string    string     string
  ─────── ───────────────────────── ───────── ──────────
0      id                     int64       8 B       0 MB
1    hood StringDType(na_object='')      16 B       1 MB
2 zipcode StringDType(na_object='')      16 B       1 MB
3  guests                     int64       8 B       0 MB
4    sqft                   float64       8 B       0 MB
5   price                     int64       8 B       0 MB
6   TOTAL                        --      64 B       3 MB
.
None
print_na_counts()[source]

Print counts of missing values by column.

>>> data = di.read_csv("data/listings.csv")
>>> data.print_na_counts()
.
  COLUMN     NNA    PNA
  string float64 string
  ────── ─────── ──────
0   sqft   49134  99.2%
.
None
rbind(*others)[source]

Return data frame with rows from others added.

>>> data = di.read_csv("data/listings.csv")
>>> data.rbind(data)
.
     id      hood zipcode guests    sqft price
  int64    string  string  int64 float64 int64
  ───── ───────── ─────── ────── ─────── ─────
0  2060 Manhattan   10040      2     nan   100
1  2595 Manhattan   10018      2     nan   225
2  3831  Brooklyn   11238      3     500    89
3  5099 Manhattan   10016      2     nan   200
4  5121  Brooklyn   11216      2     nan    60
5  5136  Brooklyn   11232      4     nan   253
6  5178 Manhattan   10019      2     nan    79
7  5203 Manhattan   10025      1     nan    79
8  5238 Manhattan   10002      2     nan   150
9  5441 Manhattan   10036      2     nan    99
.
... 99060 rows total
classmethod read_csv(path, *, encoding='utf-8', sep=',', header=True, columns=[], dtypes={})[source]

Return a new data frame from CSV file path.

Will automatically decompress if path ends in .bz2|.gz|.xz. columns is an optional list of columns to limit to. dtypes is an optional dict mapping column names to NumPy datatypes.

classmethod read_json(path, *, encoding='utf-8', columns=[], dtypes={}, **kwargs)[source]

Return a new data frame from JSON file path.

Will automatically decompress if path ends in .bz2|.gz|.xz. columns is an optional list of columns to limit to. dtypes is an optional dict mapping column names to NumPy datatypes. kwargs are passed to json.load.

classmethod read_npz(path, *, allow_pickle=True)[source]

Return a new data frame from NumPy file path.

See numpy.load for an explanation of allow_pickle: https://numpy.org/doc/stable/reference/generated/numpy.load.html

classmethod read_parquet(path, *, columns=[], dtypes={})[source]

Return a new data frame from Parquet file path.

columns is an optional list of columns to limit to. dtypes is an optional dict mapping column names to NumPy datatypes.

classmethod read_pickle(path)[source]

Return a new data frame from Pickle file path.

Will automatically decompress if path ends in .bz2|.gz|.xz.

rename(**to_from_pairs)[source]

Return data frame with columns renamed.

>>> data = di.read_csv("data/listings.csv")
>>> data.rename(listing_id="id")
.
  listing_id      hood zipcode guests    sqft price
       int64    string  string  int64 float64 int64
  ────────── ───────── ─────── ────── ─────── ─────
0       2060 Manhattan   10040      2     nan   100
1       2595 Manhattan   10018      2     nan   225
2       3831  Brooklyn   11238      3     500    89
3       5099 Manhattan   10016      2     nan   200
4       5121  Brooklyn   11216      2     nan    60
5       5136  Brooklyn   11232      4     nan   253
6       5178 Manhattan   10019      2     nan    79
7       5203 Manhattan   10025      1     nan    79
8       5238 Manhattan   10002      2     nan   150
9       5441 Manhattan   10036      2     nan    99
.
... 49530 rows total
sample(n=None)[source]

Return randomly chosen n rows.

>>> data = di.read_csv("data/listings.csv")
>>> data.sample(5)
.
        id      hood zipcode guests    sqft price
     int64    string  string  int64 float64 int64
  ──────── ───────── ─────── ────── ─────── ─────
0 12611122  Brooklyn   11217      3     nan   150
1 17713865 Manhattan   10021      3     nan   150
2 22189962    Queens   11418      8     nan    80
3 35039554    Queens   11423      8       0   430
4 38483932  Brooklyn   11211      6     nan    75
.
select(*colnames)[source]

Return data frame, keeping only colnames.

>>> data = di.read_csv("data/listings.csv")
>>> data.select("id", "hood", "zipcode")
.
     id      hood zipcode
  int64    string  string
  ───── ───────── ───────
0  2060 Manhattan   10040
1  2595 Manhattan   10018
2  3831  Brooklyn   11238
3  5099 Manhattan   10016
4  5121  Brooklyn   11216
5  5136  Brooklyn   11232
6  5178 Manhattan   10019
7  5203 Manhattan   10025
8  5238 Manhattan   10002
9  5441 Manhattan   10036
.
... 49530 rows total
semi_join(other, *by)[source]

Return rows with matches in other.

by are column names, by which to look for matching rows, or tuples of column names if the correspoding column name differs between self and other.

>>> # All listings that have reviews
>>> listings = di.read_csv("data/listings.csv")
>>> reviews = di.read_csv("data/listings-reviews.csv")
>>> listings.semi_join(reviews, "id")
.
     id      hood zipcode guests    sqft price
  int64    string  string  int64 float64 int64
  ───── ───────── ─────── ────── ─────── ─────
0  2595 Manhattan   10018      2     nan   225
1  3831  Brooklyn   11238      3     500    89
2  5099 Manhattan   10016      2     nan   200
3  5121  Brooklyn   11216      2     nan    60
4  5178 Manhattan   10019      2     nan    79
5  5203 Manhattan   10025      1     nan    79
6  5238 Manhattan   10002      2     nan   150
7  5441 Manhattan   10036      2     nan    99
8  5552 Manhattan   10014      2     nan   160
9  5803  Brooklyn   11215      2     nan    89
.
... 19019 rows total
slice(rows=None, cols=None)[source]

Return a row-wise and/or column-wise subset of data frame.

Both rows and cols should be integer vectors correspoding to the indices of the rows or columns to keep.

>>> data = di.read_csv("data/listings.csv")
>>> data.slice(rows=[0, 1, 2])
.
     id      hood zipcode guests    sqft price
  int64    string  string  int64 float64 int64
  ───── ───────── ─────── ────── ─────── ─────
0  2060 Manhattan   10040      2     nan   100
1  2595 Manhattan   10018      2     nan   225
2  3831  Brooklyn   11238      3     500    89
.
>>> data.slice(cols=[0, 1, 2])
.
     id      hood zipcode
  int64    string  string
  ───── ───────── ───────
0  2060 Manhattan   10040
1  2595 Manhattan   10018
2  3831  Brooklyn   11238
3  5099 Manhattan   10016
4  5121  Brooklyn   11216
5  5136  Brooklyn   11232
6  5178 Manhattan   10019
7  5203 Manhattan   10025
8  5238 Manhattan   10002
9  5441 Manhattan   10036
.
... 49530 rows total
>>> data.slice(rows=[0, 1, 2], cols=[0, 1, 2])
.
     id      hood zipcode
  int64    string  string
  ───── ───────── ───────
0  2060 Manhattan   10040
1  2595 Manhattan   10018
2  3831  Brooklyn   11238
.
slice_off(rows=None, cols=None)[source]

Return a row-wise and/or column-wise negative subset of data frame.

Both rows and cols should be integer vectors correspoding to the indices of the rows or columns to drop.

>>> data = di.read_csv("data/listings.csv")
>>> data.slice_off(rows=[0, 1, 2])
.
     id      hood zipcode guests    sqft price
  int64    string  string  int64 float64 int64
  ───── ───────── ─────── ────── ─────── ─────
0  5099 Manhattan   10016      2     nan   200
1  5121  Brooklyn   11216      2     nan    60
2  5136  Brooklyn   11232      4     nan   253
3  5178 Manhattan   10019      2     nan    79
4  5203 Manhattan   10025      1     nan    79
5  5238 Manhattan   10002      2     nan   150
6  5441 Manhattan   10036      2     nan    99
7  5552 Manhattan   10014      2     nan   160
8  5803  Brooklyn   11215      2     nan    89
9  6021 Manhattan   10025      1     nan    85
.
... 49527 rows total
>>> data.slice_off(cols=[0, 1, 2])
.
  guests    sqft price
   int64 float64 int64
  ────── ─────── ─────
0      2     nan   100
1      2     nan   225
2      3     500    89
3      2     nan   200
4      2     nan    60
5      4     nan   253
6      2     nan    79
7      1     nan    79
8      2     nan   150
9      2     nan    99
.
... 49530 rows total
>>> data.slice_off(rows=[0, 1, 2], cols=[0, 1, 2])
.
  guests    sqft price
   int64 float64 int64
  ────── ─────── ─────
0      2     nan   200
1      2     nan    60
2      4     nan   253
3      2     nan    79
4      1     nan    79
5      2     nan   150
6      2     nan    99
7      2     nan   160
8      2     nan    89
9      1     nan    85
.
... 49527 rows total
sort(**colname_dir_pairs)[source]

Return rows in sorted order.

colname_dir_pairs defines the sort order by column name with dir being 1 for ascending sort, -1 for descending.

>>> data = di.read_csv("data/listings.csv")
>>> data.sort(hood=1, zipcode=1)
.
       id   hood zipcode guests    sqft price
    int64 string  string  int64 float64 int64
  ─────── ────── ─────── ────── ─────── ─────
0 1026683  Bronx   10451      2     nan    85
1 3312276  Bronx   10451      2     nan    75
2 3627326  Bronx   10451      2     nan   115
3 3628328  Bronx   10451      6     nan    87
4 3939086  Bronx   10451      1     nan    45
5 6180121  Bronx   10451      2     nan    69
6 6395433  Bronx   10451      4     nan    65
7 7561686  Bronx   10451      2     nan   110
8 8673141  Bronx   10451      3     nan   225
9 9312190  Bronx   10451      7     nan   159
.
... 49530 rows total
split(*by)[source]

Split data frame into groups and return a list of their rows.

>>> data = di.DataFrame(x=[1, 2, 2, 3, 3, 3])
>>> data.split("x")
[[ 0 ] int64, [ 1 2 ] int64, [ 3 4 5 ] int64]
tail(n=None)[source]

Return the last n rows.

>>> data = di.read_csv("data/listings.csv")
>>> data.tail(5)
.
        id      hood zipcode guests    sqft price
     int64    string  string  int64 float64 int64
  ──────── ───────── ─────── ────── ─────── ─────
0 43702714 Manhattan   10003      4     nan   173
1 43702765  Brooklyn   11237      3     nan    99
2 43703128 Manhattan   10027      2     nan    49
3 43703156 Manhattan   10009      4     nan    94
4 43703359 Manhattan   10001      4     nan   249
.
to_arrow()[source]

Return data frame converted to a pyarrow.Table.

>>> data = di.read_csv("data/listings.csv")
>>> data.to_arrow()
pyarrow.Table
id: int64
hood: string
zipcode: string
guests: int64
sqft: double
price: int64
----
id: [[2060,2595,3831,5099,5121,...,43702714,43702765,43703128,43703156,43703359]]
hood: [["Manhattan","Manhattan","Brooklyn","Manhattan","Brooklyn",...,"Manhattan","Brooklyn","Manhattan","Manhattan","Manhattan"]]
zipcode: [["10040","10018","11238","10016","11216",...,"10003","11237","10027","10009","10001"]]
guests: [[2,2,3,2,2,...,4,3,2,4,4]]
sqft: [[null,null,500,null,null,...,null,null,null,null,null]]
price: [[100,225,89,200,60,...,173,99,49,94,249]]
to_json(**kwargs)[source]

Return data frame converted to a JSON string.

kwargs are passed to json.dump.

>>> data = di.read_csv("data/listings.csv")
>>> data.to_json()[:100]
[
  {
    "id": 2060,
    "hood": "Manhattan",
    "zipcode": "10040",
    "guests": 2,
    "sqft": 
to_list_of_dicts()[source]

Return data frame converted to a ListOfDicts.

>>> data = di.read_csv("data/listings.csv")
>>> data.to_list_of_dicts()
[
  {
    "id": 2060,
    "hood": "Manhattan",
    "zipcode": "10040",
    "guests": 2,
    "sqft": null,
    "price": 100
  },
  {
    "id": 2595,
    "hood": "Manhattan",
    "zipcode": "10018",
    "guests": 2,
    "sqft": null,
    "price": 225
  },
  {
    "id": 3831,
    "hood": "Brooklyn",
    "zipcode": "11238",
    "guests": 3,
    "sqft": 500.0,
    "price": 89
  }
] ... 49530 items total
to_pandas()[source]

Return data frame converted to a pandas.DataFrame.

>>> data = di.read_csv("data/listings.csv")
>>> data.to_pandas()
             id       hood zipcode  guests   sqft  price
0          2060  Manhattan   10040       2    NaN    100
1          2595  Manhattan   10018       2    NaN    225
2          3831   Brooklyn   11238       3  500.0     89
3          5099  Manhattan   10016       2    NaN    200
4          5121   Brooklyn   11216       2    NaN     60
...         ...        ...     ...     ...    ...    ...
49525  43702714  Manhattan   10003       4    NaN    173
49526  43702765   Brooklyn   11237       3    NaN     99
49527  43703128  Manhattan   10027       2    NaN     49
49528  43703156  Manhattan   10009       4    NaN     94
49529  43703359  Manhattan   10001       4    NaN    249
.
[49530 rows x 6 columns]
to_string(*, max_rows=None, max_width=None, truncate_width=None)[source]

Return data frame as a string formatted for display.

>>> data = di.read_csv("data/listings.csv")
>>> data.to_string()
.
     id      hood zipcode guests    sqft price
  int64    string  string  int64 float64 int64
  ───── ───────── ─────── ────── ─────── ─────
0  2060 Manhattan   10040      2     nan   100
1  2595 Manhattan   10018      2     nan   225
2  3831  Brooklyn   11238      3     500    89
3  5099 Manhattan   10016      2     nan   200
4  5121  Brooklyn   11216      2     nan    60
5  5136  Brooklyn   11232      4     nan   253
6  5178 Manhattan   10019      2     nan    79
7  5203 Manhattan   10025      1     nan    79
8  5238 Manhattan   10002      2     nan   150
9  5441 Manhattan   10036      2     nan    99
.
... 49530 rows total
unique(*colnames)[source]

Return unique rows by colnames.

>>> data = di.read_csv("data/listings.csv")
>>> data.unique("hood")
.
     id          hood zipcode guests    sqft price
  int64        string  string  int64 float64 int64
  ───── ───────────── ─────── ────── ─────── ─────
0  2060     Manhattan   10040      2     nan   100
1  3831      Brooklyn   11238      3     500    89
2 12937        Queens   11101      4     nan   130
3 42882 Staten Island   10301      2     nan    70
4 44096         Bronx   10452      1       9    38
.
unselect(*colnames)[source]

Return data frame, dropping colnames.

>>> data = di.read_csv("data/listings.csv")
>>> data.unselect("guests", "sqft", "price")
.
     id      hood zipcode
  int64    string  string
  ───── ───────── ───────
0  2060 Manhattan   10040
1  2595 Manhattan   10018
2  3831  Brooklyn   11238
3  5099 Manhattan   10016
4  5121  Brooklyn   11216
5  5136  Brooklyn   11232
6  5178 Manhattan   10019
7  5203 Manhattan   10025
8  5238 Manhattan   10002
9  5441 Manhattan   10036
.
... 49530 rows total
update(other)[source]

Return data frame with columns from other added.

>>> data = di.read_csv("data/listings.csv")
>>> data.update(di.DataFrame(x=1))
.
     id      hood zipcode guests    sqft price     x
  int64    string  string  int64 float64 int64 int64
  ───── ───────── ─────── ────── ─────── ───── ─────
0  2060 Manhattan   10040      2     nan   100     1
1  2595 Manhattan   10018      2     nan   225     1
2  3831  Brooklyn   11238      3     500    89     1
3  5099 Manhattan   10016      2     nan   200     1
4  5121  Brooklyn   11216      2     nan    60     1
5  5136  Brooklyn   11232      4     nan   253     1
6  5178 Manhattan   10019      2     nan    79     1
7  5203 Manhattan   10025      1     nan    79     1
8  5238 Manhattan   10002      2     nan   150     1
9  5441 Manhattan   10036      2     nan    99     1
.
... 49530 rows total
write_csv(path, *, encoding='utf-8', header=True, sep=',')[source]

Write data frame to CSV file path.

Will automatically compress if path ends in .bz2|.gz|.xz.

write_json(path, *, encoding='utf-8', **kwargs)[source]

Write data frame to JSON file path.

Will automatically compress if path ends in .bz2|.gz|.xz. kwargs are passed to json.JSONEncoder.

write_npz(path, *, compress=False)[source]

Write data frame to NumPy file path.

write_parquet(path, **kwargs)[source]

Write data frame to Parquet file path.

kwargs are passed to pyarrow.parquet.write_table.

write_pickle(path)[source]

Write data frame to Pickle file path.

Will automatically compress if path ends in .bz2|.gz|.xz.