AI

Small models on device start winning real workloads

Latency, privacy and cost are pushing classification, extraction and routing off the cloud — and product teams are noticing users prefer it.

DO

By Daniel Okafor

AI & Infrastructure Editor · Published · 5 min read

Key takeaways

  • Classification, extraction and routing tasks are moving from cloud to on-device models
  • Latency, privacy and cost are the main reasons product teams are shifting on-device
  • Features built on-device keep working even on a poor network connection
Glowing network sphere representing a machine learning model
Glowing network sphere representing a machine learning model · Illustrative image

The interesting on-device story is not chat. It is the unglamorous middle of the product: tagging, deduplication, intent routing, redaction — tasks with tight latency budgets and no tolerance for a network round trip.

Teams shipping this way describe a second benefit that is harder to price: features keep working on a bad connection.

Why it matters

Product teams building AI features can cut latency and cost by moving specific tasks on-device rather than to the cloud. Users benefit from features that keep working without a reliable network connection.

Sources and references

  • Advisable analysis
  • Updated 14 Sept 2026 · 09:00 UTC
  • AI-assisted draft, edited and fact-checked before publication · Reviewed by Advisable Desk.

In this story

  • On-device AI

Share this story

Related stories

Daily Brief

One email each morning. The stories that actually change decisions.

Five minutes of e-commerce, AI, payments, business and startup news, edited by humans and sent before markets open.

No spam. Unsubscribe in one click.