Cutting search latency from seconds to under 100ms
One product search endpoint took 1 to 2 seconds. A second one took 5 to 20 seconds. Both ran LIKE '%term%' over a large table, which forces a full scan and, under load, a queue of unhappy shoppers.
The move to OpenSearch
OpenSearch was already in the stack, so the fix was not exotic:
- Index the searchable fields. Product name, description, and the attributes customers filter by.
- Replace the queries. The slow endpoints now hit OpenSearch instead of the relational table.
- Keep the data in sync. A SQL Server Change Data Capture pipeline pushes changes into the index in near real time.
The first endpoint dropped from 1 to 2 seconds to under 100ms. The second dropped from 5 to 20 seconds to under 80ms.
Why feature flags mattered more than the index
The most important decision was not technical. We shipped the new path behind a feature flag, with the old path still in place. That let us compare results against the legacy path, roll back by flipping a config value, and ramp traffic gradually instead of betting the storefront on a big-bang switch.
When you change something customers touch on every page, the ability to say “never mind” in seconds is worth more than the fastest query.
The lesson
Search is rarely a query problem. It is an indexing problem. If a query is slow because it scans, give it a proper index somewhere instead of adding another WITH (NOLOCK) and hoping.
Comments
Comments are not enabled yet.