# Splunk Introduces Activity-Based Pricing and Machine Data Lake to Reshape AI Data Management Costs
Splunk has unveiled a significant shift in how it charges customers for data processing, introducing activity-based pricing designed to align costs more closely with actual usage rather than sheer volume. The update arrives alongside a suite of new data architecture tools aimed at supporting the growing demands of AI-driven automation in enterprise environments.
## Activity-Based Pricing and the Machine Data Lake
The cornerstone of Splunk’s new approach is the Machine Data Lake (MDL), a bulk storage layer that sits within the broader Cisco Data Fabric framework for AI-oriented data management. Unlike traditional ingest-based pricing — which charges based on the volume of data indexed each day — or workload-based pricing, which bills for compute resources consumed during searches and analytics, the new activity-based model defers indexing until data is actually queried.
This subtle but powerful change means organizations no longer need to commit to indexing every piece of data the moment it enters the platform. Data can be stored in its raw form and only processed when it becomes relevant to an investigation or analysis, dramatically reducing the cost burden associated with large-scale telemetry collection.
“We’re trying to let organizations add a zero to the amount of data they process without seeing their bill grow proportionally,” said Kamal Hathi, Splunk’s senior vice president and general manager, speaking at a recent industry conference. The goal is to make it feasible for enterprises to retain vast quantities of operational data for AI workloads without the prohibitive costs that previously made doing so impractical.
## Why CIOs Are Taking Notice
Industry analysts believe the pricing restructuring could prompt major reconsiderations of how enterprises architect their Splunk deployments. Many organizations have spent years building elaborate data tiering strategies around ingest-based licensing, deciding carefully which logs and metrics warrant immediate indexing and which get routed to cheaper storage alternatives. By weighting search and query activity equally with ingestion in the pricing formula, the new model simplifies that calculus.
“Organizations that have been carefully curating what gets indexed versus what sits in cold storage are going to have to rethink their entire approach,” noted one analyst familiar with the changes. The shift is particularly timely as AI-driven autonomous agents become more central to operations, security, and reliability workflows, all of which demand access to large, diverse datasets.
## Federated Search Expands, New Tools Emerge
Alongside the pricing overhaul, Splunk expanded its Federated Search capability to support Amazon CloudWatch and Databricks on AWS. Federated Search allows users to query data across multiple platforms without first ingesting and indexing it into Splunk, offering another pathway to cost savings. However, latency remains a concern for real-time incident response scenarios where agents need near-instant access to data.
Splunk also introduced a Value Insights feature designed to help customers evaluate which pricing model — ingest-based, workload-based, or activity-based — delivers the greatest return for their specific data profiles. Additionally, the company confirmed that an AI-powered agent capable of optimizing storage tiering is under development within the Cisco Data Fabric framework.
## Data Architecture: From Raw Telemetry to AI-Ready Signals
Perhaps the most philosophically significant development discussed at the event was the recognition that AI automation demands a fundamentally different approach to data architecture. Rather than simply amassing ever-larger reservoirs of metrics, logs, and traces, organizations are being urged to move toward what experts call “AI-ready signals.”
This paradigm requires an additional context layer that fuses raw system telemetry with operational metadata drawn from configuration management databases, IT service management platforms, business intelligence, source code repositories, and documented runbooks. The resulting enriched dataset gives AI agents the contextual awareness needed to deliver accurate, actionable insights rather than vague or misleading outputs.
One energy company reported that early attempts at deploying AI-driven security automation without first establishing a mature asset and identity framework produced an overwhelming flood of false positives, forcing the team to start over. Once proper data curation was in place, resolution times for certain incidents dropped from over 20 minutes to under a minute — a dramatic improvement directly attributable to better-organized data.
A separate case study from a large healthcare organization highlighted similar challenges. Its initial effort to build an AI site reliability engineering agent struggled with accuracy rates as low as 10% to 15%, hampered by fragmented data sources and inconsistent integration patterns. The organization emphasized that achieving the 80% to 90% accuracy thresholds required for production use remains a work in progress, but that the data architecture improvements are yielding steady gains.
## The Cisco Connection
Splunk’s deep integration with Cisco networking assets also emerged as a competitive differentiator. A newly released Network Intelligence App consolidates Cisco network topology, device health, and event data into a unified view within Splunk, adding network context to the platform’s existing security and operational data portfolios. This combination of log management, operational analytics, and network intelligence gives Splunk a uniquely broad data foundation for AI workloads.
Analysts pointed out that while many vendors now offer some form of a “data fabric,” Splunk’s combination of decades of log management expertise within large enterprises and its native Cisco integration provides a compelling value proposition for IT buyers navigating an increasingly complex AI infrastructure landscape.
## Looking Ahead
Splunk’s announcements signal a broader industry trend toward pricing models that reward efficiency over volume, and toward data architectures purpose-built for AI rather than retrofitted from traditional observability platforms. The Machine Data Lake, activity-based pricing, and the push toward AI-ready signals collectively represent a concerted effort to make enterprise data platforms economically viable for the agentic AI era.
As one expert summarized, “The organizations that invest now in structuring their data for AI will be the ones that see the real payoff when autonomous agents become a routine part of operations.”
### Frequently Asked Questions (FAQ)
**Q: What is activity-based pricing, and how does it differ from Splunk’s previous models?**
A: Activity-based pricing charges based on when and how data is queried rather than requiring customers to index all data upon ingestion. This contrasts with ingest-based pricing (charged per volume of indexed data per day) and workload-based pricing (charged for compute resources during searches). Under the new model, data can be stored in raw form and only indexed when actually searched, resulting in significant potential savings.
**Q: What is the Machine Data Lake, and what problem does it solve?**
A: The Machine Data Lake is a bulk storage layer within Splunk’s data management framework designed to hold vast quantities of raw telemetry at low cost. It solves the problem of exorbitant storage and indexing costs that previously made it impractical for organizations to retain all their operational data for AI and analytics use cases.
**Q: How does Federated Search contribute to cost savings?**
A: Federated Search allows users to query data across multiple platforms — such as Amazon CloudWatch and Databricks on AWS — without first ingesting and indexing that data into Splunk. This eliminates unnecessary indexing charges and enables on-demand access to data stored elsewhere.
**Q: Why is data architecture so important for AI automation?**
A: AI agents require rich contextual information to produce accurate and useful outputs. Raw telemetry alone — metrics, logs, and traces — lacks the operational context needed for reliable decision-making. Enriching telemetry with metadata from configuration databases, service management tools, business context, and runbooks transforms it into “AI-ready signals” that agents can act upon with confidence.
**Q: What is the Cisco Data Fabric, and how does it relate to Splunk’s new offerings?**
A: The Cisco Data Fabric is a framework for managing AI-oriented data across an enterprise. Splunk’s Machine Data Lake, data catalog, automation features, and context layer capabilities are built within this framework, leveraging Cisco’s network data and Splunk’s log analytics expertise to create a unified data foundation for AI workloads.
**Q: What challenges remain for organizations adopting these new approaches?**
A: Key challenges include the relative novelty of the data fabric architecture, limited documentation around the context layer and vector indexing capabilities, and the ongoing effort required to curate and unify fragmented data sources. Organizations also need time to determine which pricing model delivers the best return on their specific data usage patterns.
### Conclusion
Splunk’s latest updates represent a pivotal moment for enterprise data management. By introducing activity-based pricing, expanding Federated Search, and championing a shift from raw telemetry to AI-ready signals, the company is addressing both the cost concerns and the architectural challenges that have historically hindered AI adoption at scale. The combination of these innovations, underpinned by deep Cisco networking integration, positions Splunk as a serious contender in the increasingly competitive data fabric landscape. Organizations that embrace these changes early stand to gain not only substantial cost efficiencies but also a more capable foundation for the autonomous AI agents that will define the next generation of IT operations.
Thank you for reading



