dataiter.DataFrame
__init__()
aggregate()
anti_join()
cbind()
colnames
columns
compare()
copy()
count()
deepcopy()
drop_na()
filter()
filter_out()
from_arrow()
from_json()
from_pandas()
full_join()
group_by()
head()
inner_join()
left_join()
map()
modify()
ncol
nrow
pivot_longer()
pivot_wider()
print_()
print_memory_use()
print_na_counts()
rbind()
read_csv()
read_json()
read_npz()
read_parquet()
read_pickle()
rename()
sample()
select()
semi_join()
slice()
slice_off()
sort()
split()
tail()
to_arrow()
to_json()
to_list_of_dicts()
to_pandas()
to_string()
unique()
unselect()
update()
write_csv()
write_json()
write_npz()
write_parquet()
write_pickle()
- class dataiter.DataFrame(*args, **kwargs)[source]
A class for tabular data.
DataFrame is a subclass of
dict, with columns beingDataFrameColumn, which areVector, which are NumPyndarray. This means that basicdictmethods, such asitems(),keys()andvalues()can be used iterate over and manage the data as a whole and NumPy functions and array methods can be used for fast vectorized computations on the data.Columns can be accessed by attribute notation, e.g.
data.xin addition todata["x"]. In most cases, attribute access should be more convenient and is the way recommended by dataiter. You’ll still need to use the bracket notation for any column names that are not valid identifiers, such as ones with spaces, or ones that conflict with dict methods, such as “items”.DataFrame does not support indexing directly as the bracket notation is used to refer to dict keys, i.e. columns by name. If you want to index the whole data frame object, use the method
slice(). Individual columns are indexed the same as NumPy arrays.- aggregate(**colname_function_pairs)[source]
Return group-wise calculated summaries.
Usually aggregation is preceded by grouping, which can be conveniently written via method chaining as
data.group_by(...).aggregate(...).In colname_function_pairs, function receives as an argument a data frame object, a group-wise subset of all rows. It should return a scalar value. Common aggregation functions have shorthand helpers available under
dataiter, see the guide on aggregation for details.>>> data = di.read_csv("data/listings.csv") >>> # The below aggregations are identical. Usually you'll get by >>> # with the shorthand helpers, but for complicated calculations, >>> # you might need custom lambda functions. >>> data.group_by("hood").aggregate(n=di.count(), price=di.mean("price")) . hood n price string int64 float64 ───────────── ───── ─────── 0 Bronx 1198 90.176 1 Brooklyn 19931 125.056 2 Manhattan 21963 218.855 3 Queens 6068 99.745 4 Staten Island 370 116.908 . >>> data.group_by("hood").aggregate(n=lambda x: x.nrow, price=lambda x: x.price.mean()) . hood n price string int64 float64 ───────────── ───── ─────── 0 Bronx 1198 90.176 1 Brooklyn 19931 125.056 2 Manhattan 21963 218.855 3 Queens 6068 99.745 4 Staten Island 370 116.908 .
- anti_join(other, *by)[source]
Return rows with no matches in other.
by are column names, by which to look for matching rows, or tuples of column names if the correspoding column name differs between self and other.
>>> # All listings that don't have reviews >>> listings = di.read_csv("data/listings.csv") >>> reviews = di.read_csv("data/listings-reviews.csv") >>> listings.anti_join(reviews, "id") . id hood zipcode guests sqft price int64 string string int64 float64 int64 ───── ───────── ─────── ────── ─────── ───── 0 2060 Manhattan 10040 2 nan 100 1 5136 Brooklyn 11232 4 nan 253 2 7750 Manhattan 10029 1 750 35 3 7801 Brooklyn 11211 4 nan 299 4 8700 Manhattan 10034 2 700 80 5 11943 Brooklyn 11226 1 nan 150 6 15396 Manhattan 10001 4 nan 400 7 16458 Brooklyn 11215 4 nan 225 8 20300 Manhattan 10009 2 nan 50 9 21644 Manhattan 10031 1 nan 89 . ... 30511 rows total
- cbind(*others)[source]
Return data frame with columns from others added.
>>> data = di.read_csv("data/listings.csv") >>> data.cbind(di.DataFrame(x=1)) . id hood zipcode guests sqft price x int64 string string int64 float64 int64 int64 ───── ───────── ─────── ────── ─────── ───── ───── 0 2060 Manhattan 10040 2 nan 100 1 1 2595 Manhattan 10018 2 nan 225 1 2 3831 Brooklyn 11238 3 500 89 1 3 5099 Manhattan 10016 2 nan 200 1 4 5121 Brooklyn 11216 2 nan 60 1 5 5136 Brooklyn 11232 4 nan 253 1 6 5178 Manhattan 10019 2 nan 79 1 7 5203 Manhattan 10025 1 nan 79 1 8 5238 Manhattan 10002 2 nan 150 1 9 5441 Manhattan 10036 2 nan 99 1 . ... 49530 rows total
- property colnames
Get or set column names as a list.
>>> data = di.read_csv("data/listings.csv") >>> data.head() . id hood zipcode guests sqft price int64 string string int64 float64 int64 ───── ───────── ─────── ────── ─────── ───── 0 2060 Manhattan 10040 2 nan 100 1 2595 Manhattan 10018 2 nan 225 2 3831 Brooklyn 11238 3 500 89 3 5099 Manhattan 10016 2 nan 200 4 5121 Brooklyn 11216 2 nan 60 5 5136 Brooklyn 11232 4 nan 253 6 5178 Manhattan 10019 2 nan 79 7 5203 Manhattan 10025 1 nan 79 8 5238 Manhattan 10002 2 nan 150 9 5441 Manhattan 10036 2 nan 99 . >>> data.colnames ['id', 'hood', 'zipcode', 'guests', 'sqft', 'price'] >>> data.colnames = ["a", "b", "c", "d", "e", "f"] >>> data.head() . a b c d e f int64 string string int64 float64 int64 ───── ───────── ────── ───── ─────── ───── 0 2060 Manhattan 10040 2 nan 100 1 2595 Manhattan 10018 2 nan 225 2 3831 Brooklyn 11238 3 500 89 3 5099 Manhattan 10016 2 nan 200 4 5121 Brooklyn 11216 2 nan 60 5 5136 Brooklyn 11232 4 nan 253 6 5178 Manhattan 10019 2 nan 79 7 5203 Manhattan 10025 1 nan 79 8 5238 Manhattan 10002 2 nan 150 9 5441 Manhattan 10036 2 nan 99 .
- property columns
Return columns as a list.
- compare(other, *by, ignore_columns=[], max_changed=inf)[source]
Find differences against another data frame.
by are identifier columns which are used to uniquely identify rows and match them between self and other. compare will not work if your data lacks suitable identifiers. ignore_columns is an optional list of columns, differences in which to ignore.
compare returns three data frames: added rows, removed rows and changed values. The first two are basically subsets of the rows of self and other, respectively. Changed values are returned as a data frame with one row per differing value (not per differing row). Listing changes will terminate once max_changed is reached.
Warning
compare is experimental, do not rely on it reporting all of the differences correctly. Do not try to give it two huge data frames with very little in common, unless also giving some sensible value for max_changed.
>>> old = di.read_csv("data/vehicles.csv") >>> new = old.modify(hwy=lambda x: np.minimum(100, x.hwy)) >>> added, removed, changed = new.compare(old, "id") >>> changed . id column xvalue yvalue int64 string int64 int64 ───── ────── ────── ────── 0 33640 hwy 100 109 1 33396 hwy 100 108 2 34392 hwy 100 108 3 33265 hwy 100 105 4 33905 hwy 100 105 5 33558 hwy 100 102 6 34699 hwy 100 101 7 34918 hwy 100 101 8 33307 hwy 100 105 .
- count(*colnames)[source]
Return row counts grouped by colnames.
>>> data = di.read_csv("data/listings.csv") >>> data.count("hood") . hood n string int64 ───────────── ───── 0 Bronx 1198 1 Brooklyn 19931 2 Manhattan 21963 3 Queens 6068 4 Staten Island 370 .
- drop_na(*colnames)[source]
Return data frame without rows that have missing values in colnames.
>>> data = di.read_csv("data/listings.csv") >>> data.drop_na("sqft") . id hood zipcode guests sqft price int64 string string int64 float64 int64 ───── ───────── ─────── ────── ─────── ───── 0 3831 Brooklyn 11238 3 500 89 1 6848 Brooklyn 11211 3 500 140 2 7750 Manhattan 10029 1 750 35 3 8490 Brooklyn 11216 5 800 120 4 8700 Manhattan 10034 2 700 80 5 9704 Manhattan 10027 2 900 52 6 12343 Manhattan 10027 3 1 150 7 13050 Brooklyn 11221 5 1400 120 8 16974 Manhattan 10035 8 2200 225 9 17747 Brooklyn 11238 2 1000 105 . ... 396 rows total
- filter(rows=None, **colname_value_pairs)[source]
Return rows that match condition.
Filtering can be done by either rows or colname_value_pairs. rows can be either a boolean vector or a function that receives the data frame as argument and returns a boolean vector. The latter is especially useful in a method chaining context where you don’t have direct access to the data frame in question. Alternatively, colname_value_pairs provides a shorthand to check against a fixed value. See the example below of equivalent filtering all three ways.
>>> data = di.read_csv("data/listings.csv") >>> data.filter((data.hood == "Manhattan") & (data.guests == 2)) . id hood zipcode guests sqft price int64 string string int64 float64 int64 ───── ───────── ─────── ────── ─────── ───── 0 2060 Manhattan 10040 2 nan 100 1 2595 Manhattan 10018 2 nan 225 2 5099 Manhattan 10016 2 nan 200 3 5178 Manhattan 10019 2 nan 79 4 5238 Manhattan 10002 2 nan 150 5 5441 Manhattan 10036 2 nan 99 6 5552 Manhattan 10014 2 nan 160 7 8700 Manhattan 10034 2 700 80 8 9668 Manhattan 10031 2 nan 50 9 9704 Manhattan 10027 2 900 52 . ... 10209 rows total >>> data.filter(lambda x: (x.hood == "Manhattan") & (x.guests == 2)) . id hood zipcode guests sqft price int64 string string int64 float64 int64 ───── ───────── ─────── ────── ─────── ───── 0 2060 Manhattan 10040 2 nan 100 1 2595 Manhattan 10018 2 nan 225 2 5099 Manhattan 10016 2 nan 200 3 5178 Manhattan 10019 2 nan 79 4 5238 Manhattan 10002 2 nan 150 5 5441 Manhattan 10036 2 nan 99 6 5552 Manhattan 10014 2 nan 160 7 8700 Manhattan 10034 2 700 80 8 9668 Manhattan 10031 2 nan 50 9 9704 Manhattan 10027 2 900 52 . ... 10209 rows total >>> data.filter(hood="Manhattan", guests=2) . id hood zipcode guests sqft price int64 string string int64 float64 int64 ───── ───────── ─────── ────── ─────── ───── 0 2060 Manhattan 10040 2 nan 100 1 2595 Manhattan 10018 2 nan 225 2 5099 Manhattan 10016 2 nan 200 3 5178 Manhattan 10019 2 nan 79 4 5238 Manhattan 10002 2 nan 150 5 5441 Manhattan 10036 2 nan 99 6 5552 Manhattan 10014 2 nan 160 7 8700 Manhattan 10034 2 700 80 8 9668 Manhattan 10031 2 nan 50 9 9704 Manhattan 10027 2 900 52 . ... 10209 rows total
- filter_out(rows=None, **colname_value_pairs)[source]
Return rows that don’t match condition.
Filtering can be done by either rows or colname_value_pairs. rows can be either a boolean vector or a function that receives the data frame as argument and returns a boolean vector. The latter is especially useful in a method chaining context where you don’t have direct access to the data frame in question. Alternatively, colname_value_pairs provides a shorthand to check against a fixed value. See the example below of equivalent filtering all three ways.
>>> data = di.read_csv("data/listings.csv") >>> data.filter_out(data.hood == "Manhattan") . id hood zipcode guests sqft price int64 string string int64 float64 int64 ───── ──────── ─────── ────── ─────── ───── 0 3831 Brooklyn 11238 3 500 89 1 5121 Brooklyn 11216 2 nan 60 2 5136 Brooklyn 11232 4 nan 253 3 5803 Brooklyn 11215 2 nan 89 4 6848 Brooklyn 11211 3 500 140 5 7097 Brooklyn 11205 4 nan 199 6 7801 Brooklyn 11211 4 nan 299 7 8490 Brooklyn 11216 5 800 120 8 10452 Brooklyn 11238 3 nan 70 9 10962 Brooklyn 11215 2 nan 89 . ... 27567 rows total >>> data.filter_out(lambda x: x.hood == "Manhattan") . id hood zipcode guests sqft price int64 string string int64 float64 int64 ───── ──────── ─────── ────── ─────── ───── 0 3831 Brooklyn 11238 3 500 89 1 5121 Brooklyn 11216 2 nan 60 2 5136 Brooklyn 11232 4 nan 253 3 5803 Brooklyn 11215 2 nan 89 4 6848 Brooklyn 11211 3 500 140 5 7097 Brooklyn 11205 4 nan 199 6 7801 Brooklyn 11211 4 nan 299 7 8490 Brooklyn 11216 5 800 120 8 10452 Brooklyn 11238 3 nan 70 9 10962 Brooklyn 11215 2 nan 89 . ... 27567 rows total >>> data.filter_out(hood="Manhattan") . id hood zipcode guests sqft price int64 string string int64 float64 int64 ───── ──────── ─────── ────── ─────── ───── 0 3831 Brooklyn 11238 3 500 89 1 5121 Brooklyn 11216 2 nan 60 2 5136 Brooklyn 11232 4 nan 253 3 5803 Brooklyn 11215 2 nan 89 4 6848 Brooklyn 11211 3 500 140 5 7097 Brooklyn 11205 4 nan 199 6 7801 Brooklyn 11211 4 nan 299 7 8490 Brooklyn 11216 5 800 120 8 10452 Brooklyn 11238 3 nan 70 9 10962 Brooklyn 11215 2 nan 89 . ... 27567 rows total
- classmethod from_arrow(data, *, dtypes={})[source]
Return a new data frame from
pyarrow.Tabledata.dtypes is an optional dict mapping column names to NumPy datatypes.
- classmethod from_json(string, *, columns=[], dtypes={}, **kwargs)[source]
Return a new data frame from JSON string.
columns is an optional list of columns to limit to. dtypes is an optional dict mapping column names to NumPy datatypes. kwargs are passed to
json.load.
- classmethod from_pandas(data, *, dtypes={})[source]
Return a new data frame from
pandas.DataFramedata.dtypes is an optional dict mapping column names to NumPy datatypes.
- full_join(other, *by)[source]
Return data frame with matching rows merged from self and other.
full_join keeps all rows from both data frames, merging matching ones. If there are multiple matches, the first one will be used. For rows, for which matches are not found, missing values are added.
by are column names, by which to look for matching rows, or tuples of column names if the correspoding column name differs between self and other.
>>> listings = di.read_csv("data/listings.csv") >>> reviews = di.read_csv("data/listings-reviews.csv") >>> listings.full_join(reviews, "id") . id hood zipcode guests sqft price reviews rating int64 string string int64 float64 int64 float64 float64 ───── ───────── ─────── ────── ─────── ───── ─────── ─────── 0 2060 Manhattan 10040 2 nan 100 nan nan 1 2595 Manhattan 10018 2 nan 225 48 94 2 3831 Brooklyn 11238 3 500 89 322 89 3 5099 Manhattan 10016 2 nan 200 78 90 4 5121 Brooklyn 11216 2 nan 60 50 90 5 5136 Brooklyn 11232 4 nan 253 nan nan 6 5178 Manhattan 10019 2 nan 79 473 84 7 5203 Manhattan 10025 1 nan 79 118 98 8 5238 Manhattan 10002 2 nan 150 161 94 9 5441 Manhattan 10036 2 nan 99 213 97 . ... 49530 rows total
- group_by(*colnames)[source]
Return data frame with colnames set for grouped operations, such as
aggregate().
- head(n=None)[source]
Return the first n rows.
>>> data = di.read_csv("data/listings.csv") >>> data.head(5) . id hood zipcode guests sqft price int64 string string int64 float64 int64 ───── ───────── ─────── ────── ─────── ───── 0 2060 Manhattan 10040 2 nan 100 1 2595 Manhattan 10018 2 nan 225 2 3831 Brooklyn 11238 3 500 89 3 5099 Manhattan 10016 2 nan 200 4 5121 Brooklyn 11216 2 nan 60 .
- inner_join(other, *by)[source]
Return data frame with matching rows merged from self and other.
inner_join keeps only rows found in both data frames, merging matching ones. If there are multiple matches, the first one will be used.
by are column names, by which to look for matching rows, or tuples of column names if the correspoding column name differs between self and other.
>>> listings = di.read_csv("data/listings.csv") >>> reviews = di.read_csv("data/listings-reviews.csv") >>> listings.inner_join(reviews, "id") . id hood zipcode guests sqft price reviews rating int64 string string int64 float64 int64 int64 float64 ───── ───────── ─────── ────── ─────── ───── ─────── ─────── 0 2595 Manhattan 10018 2 nan 225 48 94 1 3831 Brooklyn 11238 3 500 89 322 89 2 5099 Manhattan 10016 2 nan 200 78 90 3 5121 Brooklyn 11216 2 nan 60 50 90 4 5178 Manhattan 10019 2 nan 79 473 84 5 5203 Manhattan 10025 1 nan 79 118 98 6 5238 Manhattan 10002 2 nan 150 161 94 7 5441 Manhattan 10036 2 nan 99 213 97 8 5552 Manhattan 10014 2 nan 160 66 97 9 5803 Brooklyn 11215 2 nan 89 180 94 . ... 19019 rows total
- left_join(other, *by)[source]
Return data frame with matching rows merged from self and other.
left_join keeps all rows in self, merging matching ones. If there are multiple matches, the first one will be used. For rows, for which matches are not found, missing values are added.
by are column names, by which to look for matching rows, or tuples of column names if the correspoding column name differs between self and other.
>>> listings = di.read_csv("data/listings.csv") >>> reviews = di.read_csv("data/listings-reviews.csv") >>> listings.left_join(reviews, "id") . id hood zipcode guests sqft price reviews rating int64 string string int64 float64 int64 float64 float64 ───── ───────── ─────── ────── ─────── ───── ─────── ─────── 0 2060 Manhattan 10040 2 nan 100 nan nan 1 2595 Manhattan 10018 2 nan 225 48 94 2 3831 Brooklyn 11238 3 500 89 322 89 3 5099 Manhattan 10016 2 nan 200 78 90 4 5121 Brooklyn 11216 2 nan 60 50 90 5 5136 Brooklyn 11232 4 nan 253 nan nan 6 5178 Manhattan 10019 2 nan 79 473 84 7 5203 Manhattan 10025 1 nan 79 118 98 8 5238 Manhattan 10002 2 nan 150 161 94 9 5441 Manhattan 10036 2 nan 99 213 97 . ... 49530 rows total
- map(function)[source]
Apply function to each row in data.
function receives as arguments the full data frame and the loop index. The return value will be a list of whatever function returns.
Note that map is an inefficient method as it iterates over rows instead of doing vectorized computation. map is mostly intended for complicated conditional cases that are difficult to express in vectorized form.
>>> data = di.read_csv("data/listings-reviews.csv") >>> data.map(lambda x, i: (x.reviews[i], x.rating[i]))[:3] [(np.int64(48), np.float64(94.0)), (np.int64(322), np.float64(89.0)), (np.int64(78), np.float64(90.0))]
- modify(**colname_value_pairs)[source]
Return data frame with columns modified.
In colname_value_pairs, value can be either a vector or a function that receives the data frame as argument and returns a vector. See the example below of equivalent modification with both ways.
Note that column modification can often be done simpler with a plain assignment, such as
data.price_per_guest = data.price / data.guests. modify just allows you to do the same in a method chain context.>>> data = di.read_csv("data/listings.csv") >>> data.modify(price_per_guest=data.price/data.guests) . id hood zipcode guests sqft price price_per_guest int64 string string int64 float64 int64 float64 ───── ───────── ─────── ────── ─────── ───── ─────────────── 0 2060 Manhattan 10040 2 nan 100 50.000 1 2595 Manhattan 10018 2 nan 225 112.500 2 3831 Brooklyn 11238 3 500 89 29.667 3 5099 Manhattan 10016 2 nan 200 100.000 4 5121 Brooklyn 11216 2 nan 60 30.000 5 5136 Brooklyn 11232 4 nan 253 63.250 6 5178 Manhattan 10019 2 nan 79 39.500 7 5203 Manhattan 10025 1 nan 79 79.000 8 5238 Manhattan 10002 2 nan 150 75.000 9 5441 Manhattan 10036 2 nan 99 49.500 . ... 49530 rows total >>> data.modify(price_per_guest=lambda x: x.price / x.guests) . id hood zipcode guests sqft price price_per_guest int64 string string int64 float64 int64 float64 ───── ───────── ─────── ────── ─────── ───── ─────────────── 0 2060 Manhattan 10040 2 nan 100 50.000 1 2595 Manhattan 10018 2 nan 225 112.500 2 3831 Brooklyn 11238 3 500 89 29.667 3 5099 Manhattan 10016 2 nan 200 100.000 4 5121 Brooklyn 11216 2 nan 60 30.000 5 5136 Brooklyn 11232 4 nan 253 63.250 6 5178 Manhattan 10019 2 nan 79 39.500 7 5203 Manhattan 10025 1 nan 79 79.000 8 5238 Manhattan 10002 2 nan 150 75.000 9 5441 Manhattan 10036 2 nan 99 49.500 . ... 49530 rows total
If the data frame is grouped, then colname_value_pairs need to be functions, which are applied to group-wise subsets of the data frame. A common use for this is calculating group-wise fractions.
>>> data = di.DataFrame(g=[1, 2, 2, 3, 3, 3]) >>> data.group_by("g").modify(f=lambda x: 1 / x.nrow) . g f int64 float64 ───── ─────── 0 1 1.00000 1 2 0.50000 2 2 0.50000 3 3 0.33333 4 3 0.33333 5 3 0.33333 .
- property ncol
Return the amount of columns.
>>> data = di.read_csv("data/listings.csv") >>> data.ncol 6
- property nrow
Return the amount of rows.
>>> data = di.read_csv("data/listings.csv") >>> data.nrow 49530
- pivot_longer(*, ids=None, names=None, values=None)[source]
Pivot data from wide to long format.
ids should be the name or a list of names of identifier columns. All other columns are considered variable columns and will be pivoted. names and values are names of columns into which the variable names and values are put in the result.
>>> wide = di.read_csv("data/listings.csv") >>> wide . id hood zipcode guests sqft price int64 string string int64 float64 int64 ───── ───────── ─────── ────── ─────── ───── 0 2060 Manhattan 10040 2 nan 100 1 2595 Manhattan 10018 2 nan 225 2 3831 Brooklyn 11238 3 500 89 3 5099 Manhattan 10016 2 nan 200 4 5121 Brooklyn 11216 2 nan 60 5 5136 Brooklyn 11232 4 nan 253 6 5178 Manhattan 10019 2 nan 79 7 5203 Manhattan 10025 1 nan 79 8 5238 Manhattan 10002 2 nan 150 9 5441 Manhattan 10036 2 nan 99 . ... 49530 rows total >>> wide.pivot_longer(ids="id", names="name", values="value") . id name value float64 string object ─────── ─────── ───────── 0 2060 hood Manhattan 1 2060 zipcode 10040 2 2060 guests 2 3 2060 sqft None 4 2060 price 100 5 2595 hood Manhattan 6 2595 zipcode 10018 7 2595 guests 2 8 2595 sqft None 9 2595 price 225 . ... 247650 rows total
- pivot_wider(*, ids=None, names=None, values=None, rename=None)[source]
Pivot data from long to wide format.
ids should be the name or a list of names of identifier columns by which the result will be unique. If ids is not given, it defaults to all columns except names and values. names should be the name of the column that contains the names of variables, which become new columns. values should be the name of the column that contains the corresponding values. rename is an optional function that you can use to e.g. lowercase the new column names or add a prefix or suffix.
>>> long = di.read_csv("data/downloads.csv") >>> long = long.sort(date=1) >>> long . category date downloads string datetime64[D] int64 ──────── ───────────── ───────── 0 Darwin 2019-09-16 35106 1 Linux 2019-09-16 2977751 2 null 2019-09-16 70379 3 other 2019-09-16 359 4 Windows 2019-09-16 62458 5 Darwin 2019-09-17 32484 6 Linux 2019-09-17 3174634 7 null 2019-09-17 77128 8 other 2019-09-17 456 9 Windows 2019-09-17 67484 . ... 905 rows total >>> long.pivot_wider(ids="date", names="category", values="downloads", rename=str.lower) . date darwin linux null other windows datetime64[D] float64 float64 float64 float64 float64 ───────────── ─────── ─────── ─────── ─────── ─────── 0 2019-09-16 35106 2977751 70379 359 62458 1 2019-09-17 32484 3174634 77128 456 67484 2 2019-09-18 34925 3223932 77630 503 67380 3 2019-09-19 40803 3211464 77706 398 73433 4 2019-09-20 41015 3172673 71622 371 80704 5 2019-09-21 28325 2265181 40671 110 38924 6 2019-09-22 27781 2169643 41261 158 34280 7 2019-09-23 40712 3111145 72444 289 72159 8 2019-09-24 57127 3266033 78446 354 82215 9 2019-09-25 87887 3237271 76867 363 82216 . ... 181 rows total
- print_(*, max_rows=None, max_width=None, truncate_width=None)[source]
Print data frame to
sys.stdout.print_ does the same as calling Python’s builtin
printfunction, but since it’s a method, you can use it at the end of a method chain instead of wrapping aprintcall around the whole chain.>>> di.read_csv("data/listings.csv").print_() . id hood zipcode guests sqft price int64 string string int64 float64 int64 ───── ───────── ─────── ────── ─────── ───── 0 2060 Manhattan 10040 2 nan 100 1 2595 Manhattan 10018 2 nan 225 2 3831 Brooklyn 11238 3 500 89 3 5099 Manhattan 10016 2 nan 200 4 5121 Brooklyn 11216 2 nan 60 5 5136 Brooklyn 11232 4 nan 253 6 5178 Manhattan 10019 2 nan 79 7 5203 Manhattan 10025 1 nan 79 8 5238 Manhattan 10002 2 nan 150 9 5441 Manhattan 10036 2 nan 99 . ... 49530 rows total None
- print_memory_use()[source]
Print memory use by column and total.
>>> data = di.read_csv("data/listings.csv") >>> data.print_memory_use() . COLUMN DTYPE ITEM_SIZE TOTAL_SIZE string string string string ─────── ───────────────────────── ───────── ────────── 0 id int64 8 B 0 MB 1 hood StringDType(na_object='') 16 B 1 MB 2 zipcode StringDType(na_object='') 16 B 1 MB 3 guests int64 8 B 0 MB 4 sqft float64 8 B 0 MB 5 price int64 8 B 0 MB 6 TOTAL -- 64 B 3 MB . None
- print_na_counts()[source]
Print counts of missing values by column.
>>> data = di.read_csv("data/listings.csv") >>> data.print_na_counts() . COLUMN NNA PNA string float64 string ────── ─────── ────── 0 sqft 49134 99.2% . None
- rbind(*others)[source]
Return data frame with rows from others added.
>>> data = di.read_csv("data/listings.csv") >>> data.rbind(data) . id hood zipcode guests sqft price int64 string string int64 float64 int64 ───── ───────── ─────── ────── ─────── ───── 0 2060 Manhattan 10040 2 nan 100 1 2595 Manhattan 10018 2 nan 225 2 3831 Brooklyn 11238 3 500 89 3 5099 Manhattan 10016 2 nan 200 4 5121 Brooklyn 11216 2 nan 60 5 5136 Brooklyn 11232 4 nan 253 6 5178 Manhattan 10019 2 nan 79 7 5203 Manhattan 10025 1 nan 79 8 5238 Manhattan 10002 2 nan 150 9 5441 Manhattan 10036 2 nan 99 . ... 99060 rows total
- classmethod read_csv(path, *, encoding='utf-8', sep=',', header=True, columns=[], dtypes={})[source]
Return a new data frame from CSV file path.
Will automatically decompress if path ends in
.bz2|.gz|.xz. columns is an optional list of columns to limit to. dtypes is an optional dict mapping column names to NumPy datatypes.
- classmethod read_json(path, *, encoding='utf-8', columns=[], dtypes={}, **kwargs)[source]
Return a new data frame from JSON file path.
Will automatically decompress if path ends in
.bz2|.gz|.xz. columns is an optional list of columns to limit to. dtypes is an optional dict mapping column names to NumPy datatypes. kwargs are passed tojson.load.
- classmethod read_npz(path, *, allow_pickle=True)[source]
Return a new data frame from NumPy file path.
See numpy.load for an explanation of allow_pickle: https://numpy.org/doc/stable/reference/generated/numpy.load.html
- classmethod read_parquet(path, *, columns=[], dtypes={})[source]
Return a new data frame from Parquet file path.
columns is an optional list of columns to limit to. dtypes is an optional dict mapping column names to NumPy datatypes.
- classmethod read_pickle(path)[source]
Return a new data frame from Pickle file path.
Will automatically decompress if path ends in
.bz2|.gz|.xz.
- rename(**to_from_pairs)[source]
Return data frame with columns renamed.
>>> data = di.read_csv("data/listings.csv") >>> data.rename(listing_id="id") . listing_id hood zipcode guests sqft price int64 string string int64 float64 int64 ────────── ───────── ─────── ────── ─────── ───── 0 2060 Manhattan 10040 2 nan 100 1 2595 Manhattan 10018 2 nan 225 2 3831 Brooklyn 11238 3 500 89 3 5099 Manhattan 10016 2 nan 200 4 5121 Brooklyn 11216 2 nan 60 5 5136 Brooklyn 11232 4 nan 253 6 5178 Manhattan 10019 2 nan 79 7 5203 Manhattan 10025 1 nan 79 8 5238 Manhattan 10002 2 nan 150 9 5441 Manhattan 10036 2 nan 99 . ... 49530 rows total
- sample(n=None)[source]
Return randomly chosen n rows.
>>> data = di.read_csv("data/listings.csv") >>> data.sample(5) . id hood zipcode guests sqft price int64 string string int64 float64 int64 ──────── ───────── ─────── ────── ─────── ───── 0 12611122 Brooklyn 11217 3 nan 150 1 17713865 Manhattan 10021 3 nan 150 2 22189962 Queens 11418 8 nan 80 3 35039554 Queens 11423 8 0 430 4 38483932 Brooklyn 11211 6 nan 75 .
- select(*colnames)[source]
Return data frame, keeping only colnames.
>>> data = di.read_csv("data/listings.csv") >>> data.select("id", "hood", "zipcode") . id hood zipcode int64 string string ───── ───────── ─────── 0 2060 Manhattan 10040 1 2595 Manhattan 10018 2 3831 Brooklyn 11238 3 5099 Manhattan 10016 4 5121 Brooklyn 11216 5 5136 Brooklyn 11232 6 5178 Manhattan 10019 7 5203 Manhattan 10025 8 5238 Manhattan 10002 9 5441 Manhattan 10036 . ... 49530 rows total
- semi_join(other, *by)[source]
Return rows with matches in other.
by are column names, by which to look for matching rows, or tuples of column names if the correspoding column name differs between self and other.
>>> # All listings that have reviews >>> listings = di.read_csv("data/listings.csv") >>> reviews = di.read_csv("data/listings-reviews.csv") >>> listings.semi_join(reviews, "id") . id hood zipcode guests sqft price int64 string string int64 float64 int64 ───── ───────── ─────── ────── ─────── ───── 0 2595 Manhattan 10018 2 nan 225 1 3831 Brooklyn 11238 3 500 89 2 5099 Manhattan 10016 2 nan 200 3 5121 Brooklyn 11216 2 nan 60 4 5178 Manhattan 10019 2 nan 79 5 5203 Manhattan 10025 1 nan 79 6 5238 Manhattan 10002 2 nan 150 7 5441 Manhattan 10036 2 nan 99 8 5552 Manhattan 10014 2 nan 160 9 5803 Brooklyn 11215 2 nan 89 . ... 19019 rows total
- slice(rows=None, cols=None)[source]
Return a row-wise and/or column-wise subset of data frame.
Both rows and cols should be integer vectors correspoding to the indices of the rows or columns to keep.
>>> data = di.read_csv("data/listings.csv") >>> data.slice(rows=[0, 1, 2]) . id hood zipcode guests sqft price int64 string string int64 float64 int64 ───── ───────── ─────── ────── ─────── ───── 0 2060 Manhattan 10040 2 nan 100 1 2595 Manhattan 10018 2 nan 225 2 3831 Brooklyn 11238 3 500 89 . >>> data.slice(cols=[0, 1, 2]) . id hood zipcode int64 string string ───── ───────── ─────── 0 2060 Manhattan 10040 1 2595 Manhattan 10018 2 3831 Brooklyn 11238 3 5099 Manhattan 10016 4 5121 Brooklyn 11216 5 5136 Brooklyn 11232 6 5178 Manhattan 10019 7 5203 Manhattan 10025 8 5238 Manhattan 10002 9 5441 Manhattan 10036 . ... 49530 rows total >>> data.slice(rows=[0, 1, 2], cols=[0, 1, 2]) . id hood zipcode int64 string string ───── ───────── ─────── 0 2060 Manhattan 10040 1 2595 Manhattan 10018 2 3831 Brooklyn 11238 .
- slice_off(rows=None, cols=None)[source]
Return a row-wise and/or column-wise negative subset of data frame.
Both rows and cols should be integer vectors correspoding to the indices of the rows or columns to drop.
>>> data = di.read_csv("data/listings.csv") >>> data.slice_off(rows=[0, 1, 2]) . id hood zipcode guests sqft price int64 string string int64 float64 int64 ───── ───────── ─────── ────── ─────── ───── 0 5099 Manhattan 10016 2 nan 200 1 5121 Brooklyn 11216 2 nan 60 2 5136 Brooklyn 11232 4 nan 253 3 5178 Manhattan 10019 2 nan 79 4 5203 Manhattan 10025 1 nan 79 5 5238 Manhattan 10002 2 nan 150 6 5441 Manhattan 10036 2 nan 99 7 5552 Manhattan 10014 2 nan 160 8 5803 Brooklyn 11215 2 nan 89 9 6021 Manhattan 10025 1 nan 85 . ... 49527 rows total >>> data.slice_off(cols=[0, 1, 2]) . guests sqft price int64 float64 int64 ────── ─────── ───── 0 2 nan 100 1 2 nan 225 2 3 500 89 3 2 nan 200 4 2 nan 60 5 4 nan 253 6 2 nan 79 7 1 nan 79 8 2 nan 150 9 2 nan 99 . ... 49530 rows total >>> data.slice_off(rows=[0, 1, 2], cols=[0, 1, 2]) . guests sqft price int64 float64 int64 ────── ─────── ───── 0 2 nan 200 1 2 nan 60 2 4 nan 253 3 2 nan 79 4 1 nan 79 5 2 nan 150 6 2 nan 99 7 2 nan 160 8 2 nan 89 9 1 nan 85 . ... 49527 rows total
- sort(**colname_dir_pairs)[source]
Return rows in sorted order.
colname_dir_pairs defines the sort order by column name with dir being
1for ascending sort,-1for descending.>>> data = di.read_csv("data/listings.csv") >>> data.sort(hood=1, zipcode=1) . id hood zipcode guests sqft price int64 string string int64 float64 int64 ─────── ────── ─────── ────── ─────── ───── 0 1026683 Bronx 10451 2 nan 85 1 3312276 Bronx 10451 2 nan 75 2 3627326 Bronx 10451 2 nan 115 3 3628328 Bronx 10451 6 nan 87 4 3939086 Bronx 10451 1 nan 45 5 6180121 Bronx 10451 2 nan 69 6 6395433 Bronx 10451 4 nan 65 7 7561686 Bronx 10451 2 nan 110 8 8673141 Bronx 10451 3 nan 225 9 9312190 Bronx 10451 7 nan 159 . ... 49530 rows total
- split(*by)[source]
Split data frame into groups and return a list of their rows.
>>> data = di.DataFrame(x=[1, 2, 2, 3, 3, 3]) >>> data.split("x") [[ 0 ] int64, [ 1 2 ] int64, [ 3 4 5 ] int64]
- tail(n=None)[source]
Return the last n rows.
>>> data = di.read_csv("data/listings.csv") >>> data.tail(5) . id hood zipcode guests sqft price int64 string string int64 float64 int64 ──────── ───────── ─────── ────── ─────── ───── 0 43702714 Manhattan 10003 4 nan 173 1 43702765 Brooklyn 11237 3 nan 99 2 43703128 Manhattan 10027 2 nan 49 3 43703156 Manhattan 10009 4 nan 94 4 43703359 Manhattan 10001 4 nan 249 .
- to_arrow()[source]
Return data frame converted to a
pyarrow.Table.>>> data = di.read_csv("data/listings.csv") >>> data.to_arrow() pyarrow.Table id: int64 hood: string zipcode: string guests: int64 sqft: double price: int64 ---- id: [[2060,2595,3831,5099,5121,...,43702714,43702765,43703128,43703156,43703359]] hood: [["Manhattan","Manhattan","Brooklyn","Manhattan","Brooklyn",...,"Manhattan","Brooklyn","Manhattan","Manhattan","Manhattan"]] zipcode: [["10040","10018","11238","10016","11216",...,"10003","11237","10027","10009","10001"]] guests: [[2,2,3,2,2,...,4,3,2,4,4]] sqft: [[null,null,500,null,null,...,null,null,null,null,null]] price: [[100,225,89,200,60,...,173,99,49,94,249]]
- to_json(**kwargs)[source]
Return data frame converted to a JSON string.
kwargs are passed to
json.dump.>>> data = di.read_csv("data/listings.csv") >>> data.to_json()[:100] [ { "id": 2060, "hood": "Manhattan", "zipcode": "10040", "guests": 2, "sqft":
- to_list_of_dicts()[source]
Return data frame converted to a
ListOfDicts.>>> data = di.read_csv("data/listings.csv") >>> data.to_list_of_dicts() [ { "id": 2060, "hood": "Manhattan", "zipcode": "10040", "guests": 2, "sqft": null, "price": 100 }, { "id": 2595, "hood": "Manhattan", "zipcode": "10018", "guests": 2, "sqft": null, "price": 225 }, { "id": 3831, "hood": "Brooklyn", "zipcode": "11238", "guests": 3, "sqft": 500.0, "price": 89 } ] ... 49530 items total
- to_pandas()[source]
Return data frame converted to a
pandas.DataFrame.>>> data = di.read_csv("data/listings.csv") >>> data.to_pandas() id hood zipcode guests sqft price 0 2060 Manhattan 10040 2 NaN 100 1 2595 Manhattan 10018 2 NaN 225 2 3831 Brooklyn 11238 3 500.0 89 3 5099 Manhattan 10016 2 NaN 200 4 5121 Brooklyn 11216 2 NaN 60 ... ... ... ... ... ... ... 49525 43702714 Manhattan 10003 4 NaN 173 49526 43702765 Brooklyn 11237 3 NaN 99 49527 43703128 Manhattan 10027 2 NaN 49 49528 43703156 Manhattan 10009 4 NaN 94 49529 43703359 Manhattan 10001 4 NaN 249 . [49530 rows x 6 columns]
- to_string(*, max_rows=None, max_width=None, truncate_width=None)[source]
Return data frame as a string formatted for display.
>>> data = di.read_csv("data/listings.csv") >>> data.to_string() . id hood zipcode guests sqft price int64 string string int64 float64 int64 ───── ───────── ─────── ────── ─────── ───── 0 2060 Manhattan 10040 2 nan 100 1 2595 Manhattan 10018 2 nan 225 2 3831 Brooklyn 11238 3 500 89 3 5099 Manhattan 10016 2 nan 200 4 5121 Brooklyn 11216 2 nan 60 5 5136 Brooklyn 11232 4 nan 253 6 5178 Manhattan 10019 2 nan 79 7 5203 Manhattan 10025 1 nan 79 8 5238 Manhattan 10002 2 nan 150 9 5441 Manhattan 10036 2 nan 99 . ... 49530 rows total
- unique(*colnames)[source]
Return unique rows by colnames.
>>> data = di.read_csv("data/listings.csv") >>> data.unique("hood") . id hood zipcode guests sqft price int64 string string int64 float64 int64 ───── ───────────── ─────── ────── ─────── ───── 0 2060 Manhattan 10040 2 nan 100 1 3831 Brooklyn 11238 3 500 89 2 12937 Queens 11101 4 nan 130 3 42882 Staten Island 10301 2 nan 70 4 44096 Bronx 10452 1 9 38 .
- unselect(*colnames)[source]
Return data frame, dropping colnames.
>>> data = di.read_csv("data/listings.csv") >>> data.unselect("guests", "sqft", "price") . id hood zipcode int64 string string ───── ───────── ─────── 0 2060 Manhattan 10040 1 2595 Manhattan 10018 2 3831 Brooklyn 11238 3 5099 Manhattan 10016 4 5121 Brooklyn 11216 5 5136 Brooklyn 11232 6 5178 Manhattan 10019 7 5203 Manhattan 10025 8 5238 Manhattan 10002 9 5441 Manhattan 10036 . ... 49530 rows total
- update(other)[source]
Return data frame with columns from other added.
>>> data = di.read_csv("data/listings.csv") >>> data.update(di.DataFrame(x=1)) . id hood zipcode guests sqft price x int64 string string int64 float64 int64 int64 ───── ───────── ─────── ────── ─────── ───── ───── 0 2060 Manhattan 10040 2 nan 100 1 1 2595 Manhattan 10018 2 nan 225 1 2 3831 Brooklyn 11238 3 500 89 1 3 5099 Manhattan 10016 2 nan 200 1 4 5121 Brooklyn 11216 2 nan 60 1 5 5136 Brooklyn 11232 4 nan 253 1 6 5178 Manhattan 10019 2 nan 79 1 7 5203 Manhattan 10025 1 nan 79 1 8 5238 Manhattan 10002 2 nan 150 1 9 5441 Manhattan 10036 2 nan 99 1 . ... 49530 rows total
- write_csv(path, *, encoding='utf-8', header=True, sep=',')[source]
Write data frame to CSV file path.
Will automatically compress if path ends in
.bz2|.gz|.xz.
- write_json(path, *, encoding='utf-8', **kwargs)[source]
Write data frame to JSON file path.
Will automatically compress if path ends in
.bz2|.gz|.xz. kwargs are passed tojson.JSONEncoder.