dataiter.ListOfDicts
__init__()
aggregate()
anti_join()
append()
clear()
copy()
deepcopy()
drop_na()
extend()
fill_missing_keys()
filter()
filter_out()
from_json()
full_join()
group_by()
head()
inner_join()
insert()
keys()
left_join()
map()
modify()
modify_if()
pluck()
print_()
print_memory_use()
print_na_counts()
read_csv()
read_json()
read_pickle()
rename()
reverse()
sample()
select()
semi_join()
sort()
split()
tail()
to_data_frame()
to_json()
to_pandas()
to_string()
unique()
unselect()
write_csv()
write_json()
write_pickle()
- class dataiter.ListOfDicts(dicts=(), *, as_is=False)[source]
A class for data as a list of dicts.
Most of the data-modifying methods return shallow copies, that is a new list of dicts that contains the same dict objects. To avoid surprises with modifying the same dicts in different objects, list of dicts marks the previous object “obsolete” upon returning a modified copy. Any attempted operations on the obsolete object will print a warning once per object. Usually, if you see this warning, you’ll want to call
deepcopy()to create a new, completely independent object.List of dicts is a subclass of list. This means that if you need fast in-place methods instead of the regular ones that return shallow copies, you can use those from the list baseclass. A common example is appending items one by one in a for loop: instead of
data = data.append(item), you can dolist.append(data, item).Contained dicts are upon initialization converted to
attd.AttributeDict, which is a simple subclass ofdictthat provides attribute access to dict keys. This means that you can access keys as e.g.data[0].xin addition todata[0]["x"]. In most cases, attribute access should be more convenient and is the way recommended by dataiter. You’ll still need to use the bracket notation for any keys that are not valid identifiers, such as keys with spaces, or ones that conflict with dict methods, such as “items”.https://github.com/otsaloma/attd
- __init__(dicts=(), *, as_is=False)[source]
Return a new list of dicts.
dicts is the data to hold, any kind of a sequence of dicts.
as_is can be set to
Trueto not convert the dicts toattd.AttributeDict. This conversion can be skipped for a small speed gain if you know that dicts are already attribute dicts. Note that regular dicts will not work, the conversion needs to be done at some point.
- aggregate(**key_function_pairs)[source]
Return group-wise calculated summaries.
Usually aggregation is preceded by grouping, which can be conveniently written via method chaining as
data.group_by(...).aggregate(...).In key_function_pairs, function receives as an argument a list of dicts object, a group-wise subset of all items. It can return any kind of value, it will end up as-is in the output.
>>> from statistics import mean >>> data = di.read_json("data/listings.json") >>> data.group_by("hood").aggregate(n=len, price=lambda x: mean(x.pluck("price"))) [ { "hood": "Bronx", "n": 1198, "price": 90.17612687813022 }, { "hood": "Brooklyn", "n": 19931, "price": 125.05619386884753 }, { "hood": "Manhattan", "n": 21963, "price": 218.8551655056231 } ] ... 5 items total
- anti_join(other, *by)[source]
Return items with no matches in other.
by are keys, by which to look for matching items, or tuples of keys if the correspoding key differs between self and other.
>>> # All listings that don't have reviews >>> listings = di.read_json("data/listings.json") >>> reviews = di.read_json("data/listings-reviews.json") >>> listings.anti_join(reviews, "id") [ { "id": 2060, "hood": "Manhattan", "zipcode": "10040", "guests": 2, "sqft": null, "price": 100 }, { "id": 5136, "hood": "Brooklyn", "zipcode": "11232", "guests": 4, "sqft": null, "price": 253 }, { "id": 7750, "hood": "Manhattan", "zipcode": "10029", "guests": 1, "sqft": 750.0, "price": 35 } ] ... 30511 items total
- append(item)[source]
Return list with item added to the end.
>>> data = di.read_json("data/listings.json") >>> data = data.append(dict.fromkeys(data[0].keys())) >>> data.tail() [ { "id": 43703156, "hood": "Manhattan", "zipcode": "10009", "guests": 4, "sqft": null, "price": 94 }, { "id": 43703359, "hood": "Manhattan", "zipcode": "10001", "guests": 4, "sqft": null, "price": 249 }, { "id": null, "hood": null, "zipcode": null, "guests": null, "sqft": null, "price": null } ]
- clear()[source]
Return list with all items removed.
>>> data = di.read_json("data/listings.json") >>> data.clear() []
- drop_na(*keys)[source]
Return list without items that have missing values in keys.
>>> data = di.read_json("data/listings.json") >>> data.drop_na("sqft") [ { "id": 3831, "hood": "Brooklyn", "zipcode": "11238", "guests": 3, "sqft": 500.0, "price": 89 }, { "id": 6848, "hood": "Brooklyn", "zipcode": "11211", "guests": 3, "sqft": 500.0, "price": 140 }, { "id": 7750, "hood": "Manhattan", "zipcode": "10029", "guests": 1, "sqft": 750.0, "price": 35 } ] ... 396 items total
- extend(other)[source]
Return list with items from other added to the end.
>>> data = di.read_json("data/listings.json") >>> data = data.extend([dict.fromkeys(data[0].keys())]) >>> data.tail() [ { "id": 43703156, "hood": "Manhattan", "zipcode": "10009", "guests": 4, "sqft": null, "price": 94 }, { "id": 43703359, "hood": "Manhattan", "zipcode": "10001", "guests": 4, "sqft": null, "price": 249 }, { "id": null, "hood": null, "zipcode": null, "guests": null, "sqft": null, "price": null } ]
- fill_missing_keys(**key_value_pairs)[source]
Return list with missing keys added.
If key_value_pairs not given, fill all missing keys with
None.>>> data = di.read_json("data/listings.json") >>> data = data.fill_missing_keys(price=None) >>> data = data.fill_missing_keys()
- filter(function=None, **key_value_pairs)[source]
Return items that match condition.
Filtering can be done either by function, which receives an individual item as its argument and returns
TrueorFalse, or by key_value_pairs, which are a shorthand to check against a fixed value. See the example below of equivalent filtering with both ways.>>> data = di.read_json("data/listings.json") >>> data.filter(lambda x: x.hood == "Manhattan" and x.guests == 2) [ { "id": 2060, "hood": "Manhattan", "zipcode": "10040", "guests": 2, "sqft": null, "price": 100 }, { "id": 2595, "hood": "Manhattan", "zipcode": "10018", "guests": 2, "sqft": null, "price": 225 }, { "id": 5099, "hood": "Manhattan", "zipcode": "10016", "guests": 2, "sqft": null, "price": 200 } ] ... 10209 items total >>> data.filter(hood="Manhattan", guests=2) [ { "id": 2060, "hood": "Manhattan", "zipcode": "10040", "guests": 2, "sqft": null, "price": 100 }, { "id": 2595, "hood": "Manhattan", "zipcode": "10018", "guests": 2, "sqft": null, "price": 225 }, { "id": 5099, "hood": "Manhattan", "zipcode": "10016", "guests": 2, "sqft": null, "price": 200 } ] ... 10209 items total
- filter_out(function=None, **key_value_pairs)[source]
Return items that don’t match condition.
Filtering can be done either by function, which receives an individual item as its argument and returns
TrueorFalse, or by key_value_pairs, which are a shorthand to check against a fixed value. See the example below of equivalent filtering with both ways.>>> data = di.read_json("data/listings.json") >>> data.filter_out(lambda x: x.hood == "Manhattan") [ { "id": 3831, "hood": "Brooklyn", "zipcode": "11238", "guests": 3, "sqft": 500.0, "price": 89 }, { "id": 5121, "hood": "Brooklyn", "zipcode": "11216", "guests": 2, "sqft": null, "price": 60 }, { "id": 5136, "hood": "Brooklyn", "zipcode": "11232", "guests": 4, "sqft": null, "price": 253 } ] ... 27567 items total >>> data.filter_out(hood="Manhattan") [ { "id": 3831, "hood": "Brooklyn", "zipcode": "11238", "guests": 3, "sqft": 500.0, "price": 89 }, { "id": 5121, "hood": "Brooklyn", "zipcode": "11216", "guests": 2, "sqft": null, "price": 60 }, { "id": 5136, "hood": "Brooklyn", "zipcode": "11232", "guests": 4, "sqft": null, "price": 253 } ] ... 27567 items total
- classmethod from_json(string, *, keys=[], types={}, **kwargs)[source]
Return a new list of dicts from JSON string.
keys is an optional list of keys to limit to. types is an optional dict mapping keys to datatypes. kwargs are passed to
json.load.
- full_join(other, *by)[source]
Return list with matching items merged from self and other.
full_join keeps all items from both lists, merging matching ones. If there are multiple matches, the first one will be used. For items, for which matches are not found, no keys are added.
by are keys, by which to look for matching items, or tuples of keys if the correspoding key differs between self and other.
>>> listings = di.read_json("data/listings.json") >>> reviews = di.read_json("data/listings-reviews.json") >>> listings.full_join(reviews, "id") [ { "id": 2060, "hood": "Manhattan", "zipcode": "10040", "guests": 2, "sqft": null, "price": 100 }, { "id": 2595, "hood": "Manhattan", "zipcode": "10018", "guests": 2, "sqft": null, "price": 225, "reviews": 48, "rating": 94.0 }, { "id": 3831, "hood": "Brooklyn", "zipcode": "11238", "guests": 3, "sqft": 500.0, "price": 89, "reviews": 322, "rating": 89.0 } ] ... 49530 items total
- group_by(*keys)[source]
Return list with keys set for grouped operations, such as
aggregate().
- head(n=None)[source]
Return the first n items.
>>> data = di.read_json("data/listings.json") >>> data.head(3) [ { "id": 2060, "hood": "Manhattan", "zipcode": "10040", "guests": 2, "sqft": null, "price": 100 }, { "id": 2595, "hood": "Manhattan", "zipcode": "10018", "guests": 2, "sqft": null, "price": 225 }, { "id": 3831, "hood": "Brooklyn", "zipcode": "11238", "guests": 3, "sqft": 500.0, "price": 89 } ]
- inner_join(other, *by)[source]
Return list with matching items merged from self and other.
inner_join keeps only items found in both lists, merging matching ones. If there are multiple matches, the first one will be used.
by are keys, by which to look for matching items, or tuples of keys if the correspoding key differs between self and other.
>>> listings = di.read_json("data/listings.json") >>> reviews = di.read_json("data/listings-reviews.json") >>> listings.inner_join(reviews, "id") [ { "id": 2595, "hood": "Manhattan", "zipcode": "10018", "guests": 2, "sqft": null, "price": 225, "reviews": 48, "rating": 94.0 }, { "id": 3831, "hood": "Brooklyn", "zipcode": "11238", "guests": 3, "sqft": 500.0, "price": 89, "reviews": 322, "rating": 89.0 }, { "id": 5099, "hood": "Manhattan", "zipcode": "10016", "guests": 2, "sqft": null, "price": 200, "reviews": 78, "rating": 90.0 } ] ... 19019 items total
- insert(index, item)[source]
Return list with item inserted at index.
>>> data = di.read_json("data/listings.json") >>> data = data.insert(0, dict.fromkeys(data[0].keys())) >>> data.head() [ { "id": null, "hood": null, "zipcode": null, "guests": null, "sqft": null, "price": null }, { "id": 2060, "hood": "Manhattan", "zipcode": "10040", "guests": 2, "sqft": null, "price": 100 }, { "id": 2595, "hood": "Manhattan", "zipcode": "10018", "guests": 2, "sqft": null, "price": 225 } ]
- keys()[source]
Return an iterator over unique keys in all items.
>>> data = di.read_json("data/listings.json") >>> list(data.keys()) ['id', 'hood', 'zipcode', 'guests', 'sqft', 'price']
- left_join(other, *by)[source]
Return list with matching items merged from self and other.
left_join keeps all items in self, merging matching ones from other. If there are multiple matches, the first one will be used. For items, for which matches are not found, no keys are added.
by are keys, by which to look for matching items, or tuples of keys if the correspoding key differs between self and other.
>>> listings = di.read_json("data/listings.json") >>> reviews = di.read_json("data/listings-reviews.json") >>> listings.left_join(reviews, "id") [ { "id": 2060, "hood": "Manhattan", "zipcode": "10040", "guests": 2, "sqft": null, "price": 100 }, { "id": 2595, "hood": "Manhattan", "zipcode": "10018", "guests": 2, "sqft": null, "price": 225, "reviews": 48, "rating": 94.0 }, { "id": 3831, "hood": "Brooklyn", "zipcode": "11238", "guests": 3, "sqft": 500.0, "price": 89, "reviews": 322, "rating": 89.0 } ] ... 49530 items total
- map(function)[source]
Apply function to each item in list.
If function returns a dict, then the return value will be coerced to a
ListOfDictsinstance, otherwise the return value will be a list of whatever function returns.>>> data = di.read_json("data/listings.json") >>> data.map(lambda x: (x.guests, x.price))[:3] [(2, 100), (2, 225), (3, 89)]
- modify(**key_function_pairs)[source]
Return list with items modified.
In key_function_pairs, function receives as an argument an individual item.
>>> data = di.read_json("data/listings.json") >>> data.modify(price_per_guest=lambda x: x.price / x.guests) [ { "id": 2060, "hood": "Manhattan", "zipcode": "10040", "guests": 2, "sqft": null, "price": 100, "price_per_guest": 50.0 }, { "id": 2595, "hood": "Manhattan", "zipcode": "10018", "guests": 2, "sqft": null, "price": 225, "price_per_guest": 112.5 }, { "id": 3831, "hood": "Brooklyn", "zipcode": "11238", "guests": 3, "sqft": 500.0, "price": 89, "price_per_guest": 29.666666666666668 } ] ... 49530 items total
- modify_if(predicate, **key_function_pairs)[source]
Return list with items matching predicate modified.
predicate is a function that receives an individual item as argument and returns
Trueto modify orFalseto not modify.In key_function_pairs, function receives as an argument an individual item.
>>> data = di.read_json("data/listings.json") >>> data.modify_if(lambda x: x.sqft, price_per_sqft=lambda x: x.price / x.sqft) [ { "id": 2060, "hood": "Manhattan", "zipcode": "10040", "guests": 2, "sqft": null, "price": 100 }, { "id": 2595, "hood": "Manhattan", "zipcode": "10018", "guests": 2, "sqft": null, "price": 225 }, { "id": 3831, "hood": "Brooklyn", "zipcode": "11238", "guests": 3, "sqft": 500.0, "price": 89, "price_per_sqft": 0.178 } ] ... 49530 items total
- pluck(key, default=None)[source]
Return a list of the values of key in all items.
default is used for items in which key is not found.
>>> data = di.read_json("data/listings.json") >>> data.pluck("id")[:10] [2060, 2595, 3831, 5099, 5121, 5136, 5178, 5203, 5238, 5441]
- print_(*, max_items=None)[source]
Print list to
sys.stdout.print_ does the same as calling Python’s builtin
printfunction, but since it’s a method, you can use it at the end of a method chain instead of wrapping aprintcall around the whole chain.>>> di.read_json("data/listings.json").print_() [ { "id": 2060, "hood": "Manhattan", "zipcode": "10040", "guests": 2, "sqft": null, "price": 100 }, { "id": 2595, "hood": "Manhattan", "zipcode": "10018", "guests": 2, "sqft": null, "price": 225 }, { "id": 3831, "hood": "Brooklyn", "zipcode": "11238", "guests": 3, "sqft": 500.0, "price": 89 } ] ... 49530 items total None
- print_memory_use()[source]
Print memory use by key and total.
>>> data = di.read_json("data/listings.json") >>> data.print_memory_use() . KEY TYPE ITEM_SIZE TOTAL_SIZE string string string string ─────── ────── ───────── ────────── 0 id int 28 B 1 MB 1 hood str 49 B 2 MB 2 zipcode str 46 B 2 MB 3 guests int 28 B 1 MB 4 sqft float 16 B 1 MB 5 price int 28 B 1 MB 6 TOTAL -- 195 B 9 MB . None
- print_na_counts()[source]
Print counts of missing values by key.
Both keys entirely missing and keys with a value of
Noneare considered missing.>>> data = di.read_json("data/listings.json") >>> data.print_na_counts() Missing counts: ... sqft: 49134 (99.2%) None
- classmethod read_csv(path, *, encoding='utf-8', sep=',', header=True, keys=[], types={})[source]
Return a new list from CSV file path.
Will automatically decompress if path ends in
.bz2|.gz|.xz. keys is an optional list of keys to limit to. types is an optional dict mapping keys to datatypes.
- classmethod read_json(path, *, encoding='utf-8', keys=[], types={}, **kwargs)[source]
Return a new list from JSON file path.
Will automatically decompress if path ends in
.bz2|.gz|.xz. keys is an optional list of keys to limit to. types is an optional dict mapping keys to datatypes. kwargs are passed tojson.load.
- classmethod read_pickle(path)[source]
Return a new list from Pickle file path.
Will automatically decompress if path ends in
.bz2|.gz|.xz.
- rename(**to_from_pairs)[source]
Return items with keys renamed.
>>> data = di.read_json("data/listings.json") >>> data.rename(listing_id="id") [ { "listing_id": 2060, "hood": "Manhattan", "zipcode": "10040", "guests": 2, "sqft": null, "price": 100 }, { "listing_id": 2595, "hood": "Manhattan", "zipcode": "10018", "guests": 2, "sqft": null, "price": 225 }, { "listing_id": 3831, "hood": "Brooklyn", "zipcode": "11238", "guests": 3, "sqft": 500.0, "price": 89 } ] ... 49530 items total
- sample(n=None)[source]
Return randomly chosen n items.
>>> data = di.read_json("data/listings.json") >>> data.sample(3) [ { "id": 2238389, "hood": "Brooklyn", "zipcode": "11215", "guests": 7, "sqft": null, "price": 290 }, { "id": 32341313, "hood": "Queens", "zipcode": "11385", "guests": 2, "sqft": null, "price": 107 }, { "id": 40154152, "hood": "Bronx", "zipcode": "10467", "guests": 2, "sqft": null, "price": 89 } ]
- select(*keys)[source]
Return items, keeping only keys.
>>> data = di.read_json("data/listings.json") >>> data.select("id", "hood", "zipcode") [ { "id": 2060, "hood": "Manhattan", "zipcode": "10040" }, { "id": 2595, "hood": "Manhattan", "zipcode": "10018" }, { "id": 3831, "hood": "Brooklyn", "zipcode": "11238" } ] ... 49530 items total
- semi_join(other, *by)[source]
Return items with matches in other.
by are keys, by which to look for matching items, or tuples of keys if the correspoding key differs between self and other.
>>> # All listings that have reviews >>> listings = di.read_json("data/listings.json") >>> reviews = di.read_json("data/listings-reviews.json") >>> listings.semi_join(reviews, "id") [ { "id": 2595, "hood": "Manhattan", "zipcode": "10018", "guests": 2, "sqft": null, "price": 225 }, { "id": 3831, "hood": "Brooklyn", "zipcode": "11238", "guests": 3, "sqft": 500.0, "price": 89 }, { "id": 5099, "hood": "Manhattan", "zipcode": "10016", "guests": 2, "sqft": null, "price": 200 } ] ... 19019 items total
- sort(**key_dir_pairs)[source]
Return items in sorted order.
key_dir_pairs defines the sort order by key with dir being
1for ascending sort,-1for descending.>>> data = di.read_json("data/listings.json") >>> data.sort(hood=1, zipcode=1) [ { "id": 1026683, "hood": "Bronx", "zipcode": "10451", "guests": 2, "sqft": null, "price": 85 }, { "id": 3312276, "hood": "Bronx", "zipcode": "10451", "guests": 2, "sqft": null, "price": 75 }, { "id": 3627326, "hood": "Bronx", "zipcode": "10451", "guests": 2, "sqft": null, "price": 115 } ] ... 49530 items total
- split(*by)[source]
Split list into groups and return a list of their indices.
>>> data = di.ListOfDicts({"x": x} for x in [1, 2, 2, 3, 3, 3]) >>> data.split("x") [[0], [1, 2], [3, 4, 5]]
- tail(n=None)[source]
Return the last n items.
>>> data = di.read_json("data/listings.json") >>> data.tail(3) [ { "id": 43703128, "hood": "Manhattan", "zipcode": "10027", "guests": 2, "sqft": null, "price": 49 }, { "id": 43703156, "hood": "Manhattan", "zipcode": "10009", "guests": 4, "sqft": null, "price": 94 }, { "id": 43703359, "hood": "Manhattan", "zipcode": "10001", "guests": 4, "sqft": null, "price": 249 } ]
- to_data_frame()[source]
Return list converted to a
DataFrame.>>> data = di.read_json("data/listings.json") >>> data.to_data_frame() . id hood zipcode guests sqft price int64 string string int64 float64 int64 ───── ───────── ─────── ────── ─────── ───── 0 2060 Manhattan 10040 2 nan 100 1 2595 Manhattan 10018 2 nan 225 2 3831 Brooklyn 11238 3 500 89 3 5099 Manhattan 10016 2 nan 200 4 5121 Brooklyn 11216 2 nan 60 5 5136 Brooklyn 11232 4 nan 253 6 5178 Manhattan 10019 2 nan 79 7 5203 Manhattan 10025 1 nan 79 8 5238 Manhattan 10002 2 nan 150 9 5441 Manhattan 10036 2 nan 99 . ... 49530 rows total
- to_json(**kwargs)[source]
Return list converted to a JSON string.
kwargs are passed to
json.dump.>>> data = di.read_json("data/listings.json") >>> data.to_json()[:100] [ { "id": 2060, "hood": "Manhattan", "zipcode": "10040", "guests": 2, "sqft":
- to_pandas()[source]
Return list converted to a
pandas.DataFrame.>>> data = di.read_json("data/listings.json") >>> data.to_pandas() id hood zipcode guests sqft price 0 2060 Manhattan 10040 2 NaN 100 1 2595 Manhattan 10018 2 NaN 225 2 3831 Brooklyn 11238 3 500.0 89 3 5099 Manhattan 10016 2 NaN 200 4 5121 Brooklyn 11216 2 NaN 60 ... ... ... ... ... ... ... 49525 43702714 Manhattan 10003 4 NaN 173 49526 43702765 Brooklyn 11237 3 NaN 99 49527 43703128 Manhattan 10027 2 NaN 49 49528 43703156 Manhattan 10009 4 NaN 94 49529 43703359 Manhattan 10001 4 NaN 249 . [49530 rows x 6 columns]
- to_string(*, max_items=None)[source]
Return list as a string formatted for display.
>>> data = di.read_json("data/listings.json") >>> data.to_string() [ { "id": 2060, "hood": "Manhattan", "zipcode": "10040", "guests": 2, "sqft": null, "price": 100 }, { "id": 2595, "hood": "Manhattan", "zipcode": "10018", "guests": 2, "sqft": null, "price": 225 }, { "id": 3831, "hood": "Brooklyn", "zipcode": "11238", "guests": 3, "sqft": 500.0, "price": 89 } ] ... 49530 items total
- unique(*keys)[source]
Return unique items by keys.
>>> data = di.read_json("data/listings.json") >>> data.unique("hood") [ { "id": 2060, "hood": "Manhattan", "zipcode": "10040", "guests": 2, "sqft": null, "price": 100 }, { "id": 3831, "hood": "Brooklyn", "zipcode": "11238", "guests": 3, "sqft": 500.0, "price": 89 }, { "id": 12937, "hood": "Queens", "zipcode": "11101", "guests": 4, "sqft": null, "price": 130 } ] ... 5 items total
- unselect(*keys)[source]
Return items, dropping keys.
>>> data = di.read_json("data/listings.json") >>> data.unselect("guests", "sqft", "price") [ { "id": 2060, "hood": "Manhattan", "zipcode": "10040" }, { "id": 2595, "hood": "Manhattan", "zipcode": "10018" }, { "id": 3831, "hood": "Brooklyn", "zipcode": "11238" } ] ... 49530 items total
- write_csv(path, *, encoding='utf-8', header=True, sep=',')[source]
Write list to CSV file path.
Will automatically compress if path ends in
.bz2|.gz|.xz.