Guides Qdrant query throughput (QPS) scaling. Use when someone asks 'how to increase QPS', 'need more throughput', 'queries per second too low', 'batch search', 'read replicas', or 'how to handle more concurrent queries'.
Throughput scaling means handling more parallel queries per second. This is different from latency - throughput and latency are opposite tuning directions and cannot be optimized simultaneously on the same node.
High throughput favors fewer, larger segments so each query touches less overhead.
default_segment_number: 2) Maximizing throughputmemory: pinned on Qdrant 1.19 or newer, always_ram: true on 1.18 or older Quantizationoptimizer_cpu_budget to limit indexing CPUs (e.g. 2 on an 8-CPU node reserves 6 for queries)If a single node is saturated on CPU after applying the tuning above, scale horizontally with read replicas.
replication_factor: 2+ and route reads to replicas Distributed deploymentSee also Horizontal Scaling for general horizontal scaling guidance.
If it is not possible to keep all vectors in RAM, disk I/O can become the bottleneck for throughput. In this case:
io_uring on Linux (kernel 5.11+) io_uring articlecpu_count - 1, which is optimal for RAM-based search but may be too low for disk-based search. See configuration referenceskillbazaar install qdrant-scaling-qps --agent claudeSign in (free) to install skills with the CLI.
Author
@qdrant
on GitHub
Published by