TopKUnique#
- class ray.data.aggregate.TopKUnique(on: str, k: int, ignore_nulls: bool = False, alias_name: str | None = None, encode_lists: bool = False)[source]#
Bases:
VectorizedAggregateFnV2[Dict[str,List],List[Any]]Defines an exact top-k-by-frequency unique aggregation.
Counts value frequencies globally (summed across all blocks) and returns the
kmost frequent values. The ranking is global: a value that is only moderately frequent in every individual block but frequent overall is correctly preferred over a value that is frequent in a single block.Unlike
ApproximateTopK, the result is exact, no third-party dependency is required, and the output is a plain list of values (likeUnique) rather than value/count records. The price is that the accumulator holds every distinct value with its count until the final ranking, so memory grows with the number of distinct values.Ties are broken deterministically: values with equal counts are ordered by value (ascending), with nulls last.
Example
import ray from ray.data.aggregate import TopKUnique ds = ray.data.from_items([ {"word": "apple"}, {"word": "banana"}, {"word": "apple"}, {"word": "cherry"}, {"word": "apple"}, {"word": "banana"} ]) result = ds.aggregate(TopKUnique(on="word", k=2)) # result: {'topk_unique(word)': ['apple', 'banana']}
- Parameters:
on (str) – The name of the column to aggregate.
k (int) – The number of most frequent values to return.
ignore_nulls (bool) – Whether to ignore null values when counting. If
False(the default, matchingUnique), nulls are counted like any other value andNonecan appear in the result.alias_name (str | None) – Optional name for the resulting column. Defaults to
"topk_unique({on})".encode_lists (bool) – If
True, list-type column elements are flattened so that each list element is counted individually. IfFalse, entire lists are treated as single values (converted to tuples for hashability). Note that this is a top-level flatten (not a recursive flatten) operation.
Methods
Return the agg name (e.g., 'sum', 'mean', 'count').
Return the PyArrow
Fieldthis aggregator produces.