The scalability of clustering Algorithms in high-dimensional spaces


Contributors:
Repository of the University Library Linz | 2026

Clustering data reduces the information needed to understand a dataset by assigning the data points to groups, allowing users to gain insight into the overall structure of the data and identify representative data points that provide information about their respective groups. The applications of clustering are very broad, ranging from grouping users by interests and collecting similar images to summarizing log lines in large observability systems.
While this field's importance has grown significantly due to increased complexity within data and the introduction of large language embedding models, this growth is accompanied by additional computational costs necessary to implement clustering on large-scale systems. These increased costs can be attributed to factors such as the exceedingly high dimensionality of modern natural language embedding models and the need to processes increasingly large numbers of samples.
This thesis investigates the scalability of common clustering algorithms in high-dimensional spaces. In particular, it will focus on the field of natural language processing, given its rise in popularity and its versatile applications in domains such as log analytics, document organization, and information retrieval.
In this thesis, multiple clustering algorithms spanning various underlying cluster assumptions, such as partition-based, model-based, density-based, and hierarchical clustering, are systematically compared across three categories: runtime, cluster quality, and cluster stability. Experiments are conducted on a controlled synthetic dataset, as well as on three real natural language datasets transformed using a selection of modern, state-of-the-art embedding models, enabling both controlled analysis and realistic high-demand scenarios.
With these findings, this thesis provides concrete recommendations that help system architects design large-scale clustering systems while minimizing their computational costs, as well as increasing the value that the created groups provide to users of these systems.

Meet the authors

  • Josef Moritz
    Josef Moritz
    Researcher
See all publications

Get involved

We enable the best engineers and researchers to work on challenging problems and develop cutting-edge solutions ready to be applied to real-world use cases. If you are curious about the many exciting opportunities waiting for you.
Full wave bg