Hacked or down? We step in within 5 hours
Let's talk
All case studies
DevelopmentScaling2 min read

Peak hour used to break it. Now it takes a 10× spike without anyone touching it.

A growing enterprise's image recognition service buckled every time demand peaked. We built an API that scales out on its own: millions of images a day, answers in under a second, and 10x traffic spikes handled with no manual intervention.

Illustration: a stream of blank picture frames flies from the left into a big yellow lens, and comes out on the right sorted into four yellow trays, each showing a different simple shape: a circle, a triangle, a square and a face
millionsof images a day
<1 sresponse time, even at peak
10×traffic spikes, no manual intervention

Does your system dread its busiest hour? Let's talk

Quiet days prove nothing

An image recognition service looks fine while few people use it. You only find out whether it holds when dozens of applications call it at the same time.

Three ways the old one failed

The enterprise needed one central image recognition service for dozens of internal and external applications at once. The existing solution:

  • buckled at peak hours, because response times spiked under load,
  • could not grow sideways, because it could not scale horizontally,
  • made every new model a project, because adding one took weeks of engineering.

What they needed: an API that serves millions of requests a day without slowing down.

Four design decisions

We designed a fully scalable image recognition API server:

  1. Capacity follows demand. Auto-scaling infrastructure adjusts to the load.
  2. Many models, one pipeline. An optimised processing pipeline runs several recognition models on the same infrastructure at once: object detection, classification, text recognition and facial recognition.
  3. A queue takes the hit. A queue-based layer absorbs traffic spikes.
  4. Speed regardless of load. Edge caching and smart request batching keep response times under a second.

What it does now

  • Millions of images a day, with consistently sub-second response times.
  • Scales horizontally and automatically: 10x traffic spikes, no manual intervention.
  • Four kinds of recognition behind a single API.
  • A documented REST API: integration takes hours, not weeks.
  • 99.99% availability, with automatic failover and self-healing infrastructure.
  • Cost-efficient GPU use, thanks to batching and running models side by side.

The API processes more images in a single hour than a person could review in a lifetime.

What to take from it

  • Scaling is design, not an afterthought. If a system does not grow horizontally, every peak is a surprise.
  • One API, many models. Without a shared layer, every new model is weeks of work.
  • Let the system absorb the peak, not the user. That is what queues and batching are for.

Is your next peak a surprise waiting to happen?

Tell us where your system slows down. We reply within one business day. More on how we approach this: Scaling.