# dataiter.Vector

[`__init__()`](#dataiter.Vector.__init__)
[`as_boolean()`](#dataiter.Vector.as_boolean)
[`as_bytes()`](#dataiter.Vector.as_bytes)
[`as_date()`](#dataiter.Vector.as_date)
[`as_datetime()`](#dataiter.Vector.as_datetime)
[`as_float()`](#dataiter.Vector.as_float)
[`as_integer()`](#dataiter.Vector.as_integer)
[`as_object()`](#dataiter.Vector.as_object)
[`as_string()`](#dataiter.Vector.as_string)
[`concat()`](#dataiter.Vector.concat)
[`drop_na()`](#dataiter.Vector.drop_na)
[`dt`](#dataiter.Vector.dt)
[`dtype_label`](#dataiter.Vector.dtype_label)
[`equal()`](#dataiter.Vector.equal)
[`fast()`](#dataiter.Vector.fast)
[`get_memory_use()`](#dataiter.Vector.get_memory_use)
[`head()`](#dataiter.Vector.head)
[`is_boolean()`](#dataiter.Vector.is_boolean)
[`is_bytes()`](#dataiter.Vector.is_bytes)
[`is_datetime()`](#dataiter.Vector.is_datetime)
[`is_float()`](#dataiter.Vector.is_float)
[`is_integer()`](#dataiter.Vector.is_integer)
[`is_na()`](#dataiter.Vector.is_na)
[`is_number()`](#dataiter.Vector.is_number)
[`is_object()`](#dataiter.Vector.is_object)
[`is_string()`](#dataiter.Vector.is_string)
[`is_timedelta()`](#dataiter.Vector.is_timedelta)
[`length`](#dataiter.Vector.length)
[`map()`](#dataiter.Vector.map)
[`na_dtype`](#dataiter.Vector.na_dtype)
[`na_value`](#dataiter.Vector.na_value)
[`range()`](#dataiter.Vector.range)
[`rank()`](#dataiter.Vector.rank)
[`re`](#dataiter.Vector.re)
[`replace_na()`](#dataiter.Vector.replace_na)
[`sample()`](#dataiter.Vector.sample)
[`sort()`](#dataiter.Vector.sort)
[`str`](#dataiter.Vector.str)
[`tail()`](#dataiter.Vector.tail)
[`to_string()`](#dataiter.Vector.to_string)
[`tolist()`](#dataiter.Vector.tolist)
[`unique()`](#dataiter.Vector.unique)

### *class* dataiter.Vector(object, dtype=None)

A one-dimensional array.

Vector is a subclass of NumPy `ndarray`. Note that not all `ndarray`
methods have been overridden and thus by careless use of baseclass in-place
methods you might manage to twist the data into multi-dimensional or other
non-vector form, causing unexpected results.

[https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html](https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html)

#### \_\_init_\_(object, dtype=None)

Return a new vector.

object can be any one-dimensional sequence, such as a NumPy array,
Python list or tuple. Creating a vector from a NumPy array will be
fast, from other types slower as data types and special values will
need to be converted.

dtype is the NumPy-compatible data type for the vector. Providing
dtype will make creating the vector faster, otherwise the appropriate
data type will be guessed by introspecting the elements of object,
which is potentially slow, especially for large objects.

```pycon
>>> di.Vector([1, 2, 3], int)
[ 1 2 3 ] int64
```

#### as_boolean()

Return vector converted to boolean data type.

```pycon
>>> vector = di.Vector([0, 1])
>>> vector.as_boolean()
[ False True ] bool
```

#### as_bytes()

Return vector converted to bytes data type.

```pycon
>>> vector = di.Vector(["a", "b"])
>>> vector.as_bytes()
[ b'a' b'b' ] |S1
```

#### as_date()

Return vector converted to date data type.

```pycon
>>> vector = di.Vector(["2020-01-01"])
>>> vector.as_date()
[ 2020-01-01 ] datetime64[D]
```

#### as_datetime(precision='us')

Return vector converted to datetime data type.

```pycon
>>> vector = di.Vector(["2020-01-01T12:00:00"])
>>> vector.as_datetime()
[ 2020-01-01T12:00:00.000000 ] datetime64[us]
```

#### as_float()

Return vector converted to float data type.

```pycon
>>> vector = di.Vector([1, 2, 3])
>>> vector.as_float()
[ 1 2 3 ] float64
```

#### as_integer()

Return vector converted to integer data type.

```pycon
>>> vector = di.Vector([1.0, 2.0, 3.0])
>>> vector.as_integer()
[ 1 2 3 ] int64
```

#### as_object()

Return vector converted to object data type.

```pycon
>>> vector = di.Vector([1, 2, 3])
>>> vector.as_object()
[ 1 2 3 ] object
```

#### as_string()

Return vector converted to string data type.

```pycon
>>> vector = di.Vector([1, 2, 3])
>>> vector.as_string()
[ "1" "2" "3" ] string
```

#### concat(\*others)

Return vector with elements from others appended.

```pycon
>>> a = di.Vector([1, 2, 3])
>>> b = di.Vector([4, 5, 6])
>>> c = di.Vector([7, 8, 9])
>>> a.concat(b, c)
[ 1 2 3 4 5 6 7 8 9 ] int64
```

#### drop_na()

Return vector without missing values.

```pycon
>>> vector = di.Vector([1, 2, 3, None])
>>> vector.drop_na()
[ 1 2 3 ] float64
```

#### *property* dt *: DtProxy*

Proxy object for calling [`dataiter.dt`](dt.md#module-dataiter.dt) functions.

```pycon
>>> x = di.Vector(["2025-01-11"], np.datetime64)
>>> x.dt.year()
[ 2025 ] int64
>>> x.dt.month()
[ 1 ] int64
>>> x.dt.day()
[ 11 ] int64
```

#### *property* dtype_label

Return a human-readable label of vector data type.

```pycon
>>> vector = di.Vector(["abc", "def"])
>>> vector.dtype
StringDType(na_object='')
>>> vector.dtype_label
string
```

#### equal(other)

Return whether vectors are equal.

Equality is tested with `==`. As an exception, corresponding missing
values are considered equal as well.

```pycon
>>> a = di.Vector([1, 2, 3, None])
>>> b = di.Vector([1, 2, 3, None])
>>> a
[ 1 2 3 nan ] float64
>>> b
[ 1 2 3 nan ] float64
>>> a.equal(b)
True
```

#### *classmethod* fast(object, dtype=None)

Return a new vector.

Unlike [`__init__()`](#dataiter.Vector.__init__), this will **not** convert special values in
object. Use this only if you know object doesn’t contain special
values or if you know they are already of the correct type.

#### get_memory_use()

Return memory use in bytes.

```pycon
>>> vector = di.Vector(range(100))
>>> vector.get_memory_use()
800
```

#### head(n=None)

Return the first n elements.

```pycon
>>> vector = di.Vector(range(100))
>>> vector.head(10)
[ 0 1 2 3 4 5 6 7 8 9 ] int64
```

#### is_boolean()

Return whether vector data type is boolean.

#### is_bytes()

Return whether vector data type is bytes.

#### is_datetime()

Return whether vector data type is datetime.

Dates are considered datetimes as well.

#### is_float()

Return whether vector data type is float.

#### is_integer()

Return whether vector data type is integer.

#### is_na()

Return a boolean vector indicating missing data elements.

```pycon
>>> vector = di.Vector([1, 2, 3, None])
>>> vector
[ 1 2 3 nan ] float64
>>> vector.is_na()
[ False False False True ] bool
```

#### is_number()

Return whether vector data type is number.

#### is_object()

Return whether vector data type is object.

#### is_string()

Return whether vector data type is string.

#### is_timedelta()

Return whether vector data type is timedelta.

#### *property* length

Return the amount of elements.

```pycon
>>> vector = di.Vector(range(100))
>>> vector.length
100
```

#### map(function, \*args, dtype=None, \*\*kwargs)

Apply function element-wise and return a new vector.

```pycon
>>> import math
>>> vector = di.Vector(range(10))
>>> vector.map(math.pow, 2)
[ 0 1 4 9 16 25 36 49 64 81 ] float64
```

#### *property* na_dtype

Return the corresponding data type that can handle missing data.

You might need this for upcasting when missing data is first introduced.

```pycon
>>> vector = di.Vector([1, 2, 3])
>>> vector
[ 1 2 3 ] int64
>>> vector.put([2], vector.na_value)
Traceback (most recent call last):
  File "<string>", line 15, in <module>
    print(vector.put([2], vector.na_value))
          ~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^
ValueError: cannot convert float NaN to integer
>>> vector = vector.astype(vector.na_dtype)
>>> vector
[ 1 2 3 ] float64
>>> vector.put([2], vector.na_value)
None
>>> vector
[ 1 2 nan ] float64
```

#### *property* na_value

Return the corresponding value to use to represent missing data.

Dataiter is built on top of NumPy. NumPy doesn’t support a proper
missing value (“NA”), only data type specific values: `np.nan`,
`np.datetime64("NaT")` and `np.timedelta64("NaT")`. Dataiter
recommends the following values be used and internally supports them to
an extent.

| datetime   | `np.datetime64("NaT")`   |
|------------|--------------------------|
| float      | `np.nan`                 |
| integer    | `np.nan`                 |
| string     | `""`                     |
| timedelta  | `np.timedelta64("NaT")`  |
| other      | `None`                   |

Note that actually using these might require upcasting the vector.
Integer will need to be upcast to float to contain `np.nan`. Other,
such as boolean, will need to be upcast to object to contain `None`.

If you need to avoid object columns, you can also consider converting
booleans to float using [`as_float()`](#dataiter.Vector.as_float), which will give you 0.0 for
false and 1.0 for true. Depending on how you use the data, that might
work as well as an object vector of `True`, `False` and `None`.

#### range()

Return the minimum and maximum values as a two-element vector.

```pycon
>>> vector = di.Vector(range(100))
>>> vector.range()
[ 0 99 ] int64
```

#### rank(\*, method='min')

Return the order of elements in a sorted vector.

method determines how ties are resolved. **‘min’** assigns each of
equal values the same rank, the minimum of the set (also called
“competition ranking”). **‘max’** is the same, but assigning the
maximum of the set. **‘ordinal’** gives each element a distinct rank
with equal values ranked by their order in input. Ranks begin at 1.
Missing values are ranked last.

**References**

* [https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.rankdata.html](https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.rankdata.html)
* [https://www.rdocumentation.org/packages/base/topics/rank](https://www.rdocumentation.org/packages/base/topics/rank)

```pycon
>>> vector = di.Vector([3, 1, 1, 1, 2, 2])
>>> vector.rank(method="min")
[ 6 1 1 1 4 4 ] int64
>>> vector.rank(method="max")
[ 6 3 3 3 5 5 ] int64
>>> vector.rank(method="ordinal")
[ 6 1 2 3 4 5 ] int64
```

#### *property* re *: ReProxy*

Proxy object for calling [`dataiter.regex`](regex.md#module-dataiter.regex) functions.

```pycon
>>> x = di.Vector(["great", "fantastic"])
>>> x.re.sub(r"$", r"!")
[ "great!" "fantastic!" ] string
```

#### replace_na(value)

Return vector with missing values replaced with value.

```pycon
>>> vector = di.Vector([1, 2, 3, None])
>>> vector.replace_na(0)
[ 1 2 3 0 ] float64
```

#### sample(n=None)

Return randomly chosen n elements.

```pycon
>>> vector = di.Vector(range(100))
>>> vector.sample(10)
[ 10 29 30 32 33 37 67 75 81 89 ] int64
```

#### sort(\*, dir=1)

Return elements in sorted order.

dir is `1` for ascending sort, `-1` for descending. Missing
values are sorted last, regardless of dir.

```pycon
>>> vector = di.Vector([1, 2, 3, None])
>>> vector.sort(dir=1)
[ 1 2 3 nan ] float64
>>> vector.sort(dir=-1)
[ 3 2 1 nan ] float64
```

#### *property* str *: StrProxy*

Proxy object for calling `numpy.strings` functions.

[https://numpy.org/doc/stable/reference/routines.strings.html](https://numpy.org/doc/stable/reference/routines.strings.html)

```pycon
>>> x = di.Vector(["asdf", "1234"])
>>> x.str.isalpha()
[ True False ] bool
>>> x.str.str_len()
[ 4 4 ] int64
>>> x.str.upper()
[ "ASDF" "1234" ] string
```

#### tail(n=None)

Return the last n elements.

```pycon
>>> vector = di.Vector(range(100))
>>> vector.tail(10)
[ 90 91 92 93 94 95 96 97 98 99 ] int64
```

#### to_string(\*, max_elements=None)

Return vector as a string formatted for display.

```pycon
>>> vector = di.Vector([1/2, 1/3, 1/4])
>>> vector.to_string()
[ 0.500000 0.333333 0.250000 ] float64
```

#### to_strings(\*, ksep=None, quote=True, pad=False, truncate_width=inf)

Return vector as strings formatted for display.

```pycon
>>> vector = di.Vector([1/2, 1/3, 1/4])
>>> vector.to_strings()
[ "0.500000" "0.333333" "0.250000" ] string
```

#### tolist()

Return vector as a list with elements of matching Python builtin type.

Missing values are replaced with `None`.

#### unique()

Return unique elements.

```pycon
>>> vector = di.Vector([1, 1, 1, 2, 2, 3])
>>> vector.unique()
[ 1 2 3 ] int64
```
