The original test setup
We used oha to send requests to /v1/search for one brand with 10,000 products. The service, Postgres and load generator shared 14 cores. Most products were test rows inserted directly into SQL, alongside a real product. We wanted to test the index at that size, not the feed-import pipeline or a typical customer catalog.
The tests used specific embedding-session counts rather than today’s default. We didn’t record the exact setting for each row or the full machine specification. That limits how closely someone else can reproduce these results.
| Offered / served rps | p50 | p95 | p99 |
|---|---|---|---|
| 100 / 100 | 11.1 ms | 15.0 ms | 27.5 ms |
| 250 / 250 | 13.4 ms | 20.1 ms | 30.2 ms |
| 500 / 371 | 268 ms | 293 ms | 301 ms |
Why p95 and served throughput both matter
p95 means 95 percent of measured requests completed within that latency. It is not the slowest request and it is not an average. At 500 offered requests per second, the service served 371 and the tail grew sharply: that row shows saturation, not a claim that the service sustains 500 rps.
Query embedding was CPU-bound in the recorded test. More inference sessions initially improved throughput, but dividing the machine into too many sessions made each query slower. Increasing concurrency also retains more inference buffers and increases memory pressure. Tune against measured latency and memory together.
What we measured with tracking enabled
A separate local check on October 1, 2026 used 60 warmed HTTP requests returning 10 cards with exact result tracking enabled. It measured p50 13.7 ms and p95 29.6 ms. The first tracked card was 394 bytes. This path included durable event insertion and HTTP overhead.
The original load test, the retrieval-only integration spec and this tracking-enabled check are different measurements. Their workload and sample size differ, so the gap is not a controlled estimate of tracking overhead. Use the path your production integration actually runs when setting expectations.
Latency is useful only if the answer is right
The search_latency integration spec seeds 10,000 filler products and real catalog examples. Its canonical query asks for waterproof running shoes under 150 EUR in size 44. The target shares its description with a cheaper size-39 decoy and an over-budget size-44 decoy; ranking alone cannot make those variants eligible.
The spec checks that the right product appears and the decoys are excluded, and asserts retrieval p95 below 25 ms. It also exercises an exact model-number query so the lexical retrieval leg has to do useful work. This correctness/performance spec is separate from the oha HTTP load-test table.
docker compose up -d
TEST_DATABASE_URL=postgres://nxt:nxt@localhost:5433/nxt cargo test --test search_latency -- --test-threads=1What to measure on your catalog
Siroua parses exact constraints before semantic and keyword retrieval and does not call a generative language model in the search path. That removes one source of waiting, but embedding, database access, network transport and reporting still cost time.
For your deployment, record hardware, model, inference concurrency, catalog size, filter selectivity, result count, tracking mode and load-generator placement. Measure cold and warmed requests, latency tails, error rate and served throughput. Pair those numbers with variant-level correctness checks and realistic feed content.
Try search on your own catalog
Join the waitlist with your store URL and a few questions your shoppers ask. We’ll use those to check the results together.
Join the waitlist