TopKUnique#

class ray.data.aggregate.TopKUnique(on: str, k: int, ignore_nulls: bool = False, alias_name: str | None = None, encode_lists: bool = False)[source]#

Bases: VectorizedAggregateFnV2[Dict[str, List], List[Any]]

Defines an exact top-k-by-frequency unique aggregation.

Counts value frequencies globally (summed across all blocks) and returns the k most frequent values. The ranking is global: a value that is only moderately frequent in every individual block but frequent overall is correctly preferred over a value that is frequent in a single block.

Unlike ApproximateTopK, the result is exact, no third-party dependency is required, and the output is a plain list of values (like Unique) rather than value/count records. The price is that the accumulator holds every distinct value with its count until the final ranking, so memory grows with the number of distinct values.

Ties are broken deterministically: values with equal counts are ordered by value (ascending), with nulls last.

Example

import ray
from ray.data.aggregate import TopKUnique

ds = ray.data.from_items([
    {"word": "apple"}, {"word": "banana"}, {"word": "apple"},
    {"word": "cherry"}, {"word": "apple"}, {"word": "banana"}
])

result = ds.aggregate(TopKUnique(on="word", k=2))
# result: {'topk_unique(word)': ['apple', 'banana']}
Parameters:
  • on (str) – The name of the column to aggregate.

  • k (int) – The number of most frequent values to return.

  • ignore_nulls (bool) – Whether to ignore null values when counting. If False (the default, matching Unique), nulls are counted like any other value and None can appear in the result.

  • alias_name (str | None) – Optional name for the resulting column. Defaults to "topk_unique({on})".

  • encode_lists (bool) – If True, list-type column elements are flattened so that each list element is counted individually. If False, entire lists are treated as single values (converted to tuples for hashability). Note that this is a top-level flatten (not a recursive flatten) operation.

Methods

get_agg_name

Return the agg name (e.g., 'sum', 'mean', 'count').

output_field

Return the PyArrow Field this aggregator produces.