Engineering note · Measurements

How we measured 15 ms product search at 10k products

We measured 15 ms p95 at 100 requests per second on a 10,000-product catalog. Here’s what we tested, what that number means and how the newer tracking-enabled result differs. These are test results, not a production SLA.

The original test setup

We used oha to send requests to /v1/search for one brand with 10,000 products. The service, Postgres and load generator shared 14 cores. Most products were test rows inserted directly into SQL, alongside a real product. We wanted to test the index at that size, not the feed-import pipeline or a typical customer catalog.

The tests used specific embedding-session counts rather than today’s default. We didn’t record the exact setting for each row or the full machine specification. That limits how closely someone else can reproduce these results.

The original test setup
Offered / served rpsp50p95p99
100 / 10011.1 ms15.0 ms27.5 ms
250 / 25013.4 ms20.1 ms30.2 ms
500 / 371268 ms293 ms301 ms

Why p95 and served throughput both matter

p95 means 95 percent of measured requests completed within that latency. It is not the slowest request and it is not an average. At 500 offered requests per second, the service served 371 and the tail grew sharply: that row shows saturation, not a claim that the service sustains 500 rps.

Query embedding was CPU-bound in the recorded test. More inference sessions initially improved throughput, but dividing the machine into too many sessions made each query slower. Increasing concurrency also retains more inference buffers and increases memory pressure. Tune against measured latency and memory together.

What we measured with tracking enabled

A separate local check on October 1, 2026 used 60 warmed HTTP requests returning 10 cards with exact result tracking enabled. It measured p50 13.7 ms and p95 29.6 ms. The first tracked card was 394 bytes. This path included durable event insertion and HTTP overhead.

The original load test, the retrieval-only integration spec and this tracking-enabled check are different measurements. Their workload and sample size differ, so the gap is not a controlled estimate of tracking overhead. Use the path your production integration actually runs when setting expectations.

Latency is useful only if the answer is right

The search_latency integration spec seeds 10,000 filler products and real catalog examples. Its canonical query asks for waterproof running shoes under 150 EUR in size 44. The target shares its description with a cheaper size-39 decoy and an over-budget size-44 decoy; ranking alone cannot make those variants eligible.

The spec checks that the right product appears and the decoys are excluded, and asserts retrieval p95 below 25 ms. It also exercises an exact model-number query so the lexical retrieval leg has to do useful work. This correctness/performance spec is separate from the oha HTTP load-test table.

Run the database-backed retrieval spec from the repository root
docker compose up -d
TEST_DATABASE_URL=postgres://nxt:nxt@localhost:5433/nxt cargo test --test search_latency -- --test-threads=1

What to measure on your catalog

Siroua parses exact constraints before semantic and keyword retrieval and does not call a generative language model in the search path. That removes one source of waiting, but embedding, database access, network transport and reporting still cost time.

For your deployment, record hardware, model, inference concurrency, catalog size, filter selectivity, result count, tracking mode and load-generator placement. Measure cold and warmed requests, latency tails, error rate and served throughput. Pair those numbers with variant-level correctness checks and realistic feed content.

Try search on your own catalog

Join the waitlist with your store URL and a few questions your shoppers ask. We’ll use those to check the results together.

Join the waitlist