# Understanding and Optimizing AI Usage: New Tools Give Teams Visibility Into How Models Are Actually Being Used
## The Challenge of Making Sense of AI Traffic
As organizations increasingly rely on artificial intelligence for everyday work, IT and engineering teams face a growing blind spot. They can see that AI requests are increasing, but they often lack the context to understand what those requests actually represent. A spike in token consumption could mean anything from a developer running complex code reviews to a bot performing simple text formatting — and the difference matters enormously when it comes to cost, performance, and model selection.
A new set of analytics capabilities aims to change that by giving teams a deeper lens into their AI workloads, moving beyond raw counts and token tallies to reveal the nature of the work itself.
## From Usage Counts to Workload Context
The core insight behind these new analytics is straightforward: knowing how much AI traffic a team generates is not the same as knowing what that traffic is doing. Early-generation dashboards showed model names, request volumes, and token totals, but they couldn’t distinguish between a five-turn research task and a one-turn formatting request — even though both might consume similar resources.
The updated analytics layer introduces task-based classification, automatically grouping conversations into categories such as coding, research, writing, summarization, and data analysis. This transformation allows managers to see not just that their team is using AI heavily, but *how* they are using it. An engineering group might discover that most of its AI consumption goes toward debugging and code generation, while a product team might be relying on the same infrastructure primarily for research and document summarization.
## Identifying When Models Are Overqualified
One of the most immediately actionable features is the ability to detect when a high-capability model is being used for tasks that don’t require its full power. This phenomenon — sometimes called “model overkill” — occurs when simple or routine requests are routed to the same advanced reasoning models that excel at complex problem-solving.
The analytics view surfaces these patterns by correlating task category, model capability, conversation length, and cost. For example, a team might notice that their summarization requests, which typically complete in a single exchange, are consistently being sent to a large reasoning model designed for multi-step analytical work. That mismatch represents wasted spend and unnecessary latency.
Importantly, the system doesn’t dictate a replacement model. Instead, it surfaces the question: *Is the extra capability actually improving results?* Sometimes the answer is yes — a difficult debugging session genuinely benefits from a powerful model — and sometimes the answer is no, pointing to an opportunity to reroute work to a faster, more cost-effective option.
## Tracking the Full Cost of a Task
Another key dimension is conversation depth. Not all AI tasks are resolved in a single request-response cycle. Complex work often involves multiple back-and-forth exchanges: a user asks a question, the model responds, the user refines the prompt, the model adjusts, and so on until the task is complete.
The analytics track how many turns different categories of work require on average. A task that should take one turn but keeps extending across five or six is a signal that something in the workflow needs attention — whether that means improving the prompt, adjusting the model, or redesigning the agent’s behavior.
By combining turn data with token usage and cost metrics, teams get a complete picture of the resources consumed by each type of work, enabling more informed decisions about where to invest in optimization.
## Automatic Routing Based on Workload Signals
The analytical insights don’t just inform manual decisions — they can also drive automation. A new routing capability uses the same task classification and model-fit signals to automatically direct requests to the most appropriate model for the job.
Rather than requiring teams to write custom rules for every workload, the system evaluates each conversation in real time. It considers the task category, the complexity of the work, the trajectory of the conversation, and how well available models match the requirements. A simple summarization request might be routed to a lightweight, fast model, while a complex multi-step coding task would be directed to a more capable option.
The routing logic is designed to be nuanced: it does not simply send every request to the cheapest available model. Instead, it balances capability, cost, and latency to find the best fit for each individual task.
## How the Classification Pipeline Works
Under the hood, the task classification is handled by a dedicated processing service that analyzes AI traffic logs after they have been recorded. It examines the full conversation trajectory — including user prompts, model responses, tool calls, and intermediate results — to categorize the type of work being performed.
The classification happens asynchronously, meaning it takes place after the user has already received their response. This ensures that the analytics pipeline adds zero latency to the user experience. The tradeoff is that insights are not fully real-time; there is typically a short delay as logs are processed and aggregated.
The pipeline is built to integrate with existing identity systems. When AI traffic passes through an access layer that authenticates users, the analytics can associate activity with specific individuals, teams, and applications — all without exposing sensitive identity information in the prompts themselves.
## Who Benefits From These Capabilities
These features are designed for any team that routes AI traffic through a gateway and wants to understand its usage patterns more deeply. Common use cases include:
– **Engineering teams** looking to reduce AI costs by matching model capability to actual task complexity
– **Productivity-focused organizations** that want to understand how different departments use AI tools
– **Platform teams** responsible for governing AI adoption across an organization
– **Developers building AI-powered applications** who need visibility into how their tools are being used in production
The capabilities are available at no additional cost for users of the underlying AI gateway infrastructure.
—
## Frequently Asked Questions
**What kinds of tasks can the analytics classify?**
The system currently supports several broad categories including coding, research, writing, summarization, and data analysis. The categories are intentionally kept to a manageable set that is easy for teams to interpret, rather than trying to classify every possible type of work.
**Does this feature work with third-party AI tools like Claude, Codex, or OpenCode?**
Yes. When these tools are configured to route their traffic through the gateway and are authenticated through an access layer, the analytics can associate their activity with specific users and sessions automatically.
**Is there a cost to enable these analytics?**
The analytics features are available at no extra charge for teams already using the AI gateway. There are no additional pricing tiers or per-request fees for the classification and analysis pipeline.
**How quickly will I see data in the dashboard?**
Because classification is asynchronous and happens after requests are completed, there is a short processing delay. New conversations typically appear in the dashboard within approximately one day of being received. The analytics are best used for identifying patterns over time rather than monitoring live activity.
**Can I see which specific users are driving unusual usage patterns?**
Yes. The system is identity-aware and can surface anomalies at the individual user level, as well as at the team, application, and agent level. This helps teams identify unexpected spikes or out-of-control spending before it becomes a larger issue.
**Does the auto router change my existing model configuration?**
No. The router operates alongside your existing setup and can be enabled in a controlled beta environment. It does not override your models or change your prompts — it simply selects from the models already available to your application based on the detected task and model-fit signals.
**What data does the classification engine store or expose?**
The engine processes log metadata and derived categories for reporting and routing analysis. It does not store or expose raw prompts, and it does not turn the dashboard into a prompt browser. Underlying log body retention follows the standard logging configuration already in place.
—
## Conclusion
As artificial intelligence becomes more deeply embedded in everyday workflows, the ability to see *what* people are actually doing with AI — not just *how much* they are using it — becomes essential. The new analytics capabilities bridge that gap, giving teams the context they need to move from blind spending to informed optimization. By understanding which tasks demand powerful models, which are overqualified for simpler options, and where conversation patterns suggest workflow inefficiencies, organizations can make smarter decisions about model selection, routing, and cost control. These tools represent an important step toward making AI infrastructure as manageable and transparent as any other enterprise workload.
Thank you for reading



