# Cutting LLM Costs in Practice: An Apartment Search Agent Experiment
Large language models are powerful, but every call to a model costs money — and every unnecessary call wastes it. When I set out to build an agent that screens rental listings for renters, I wanted to understand what actually drives costs and where the biggest savings hide. This article walks through the experiment, the configurations I tested, and the surprising findings that emerged.
## The Agent I Built
The agent answers one yes-or-no question per listing and per renter requirement. A single listing checked against a single requirement is called a “pair,” so 500 listings and 5 requirements produce 2,500 pairs. The agent checks whether a listing matches rules like “two bedrooms in Austin under $1,500 that allows dogs.”
The listings come from a public dataset of 10,000 United States rental ads collected in late 2019. Each ad contains fields for the title, body text, amenities, bedroom count, bathroom count, monthly rent, square footage, city, state, and a pets-allowed field. From this dataset, I sampled 500 listings — 200 from Austin, 100 from Dallas, 100 from Houston, and 100 from other cities — using a fixed random seed for reproducibility.
I used two models: a strong model I call Sol, which costs $2.00 per million input tokens and $10.00 per million output tokens, and a cheaper model I call Luna, which costs $0.10 per million input tokens and $0.50 per million output tokens. Both models did not accept a temperature setting, so every run used the default behavior.
I traced every call with Weights & Biases (W&B) Weave, a tool that records each step of an application. The complete script is available alongside this article.
## Three Starting Points
My first version of the agent sent every single pair to the strong model, even when a city name or a rent figure already ruled the listing out. That approach required 2,500 model calls to find 101 real matches — a clear waste.
Instead of measuring just one starting point, I compared three:
1. **Send every pair to the model** — the naive approach, one call per pair.
2. **City-filtered** — the feed has already been searched by city, so only pairs where the city matches reach the model.
3. **Database query first** — a deterministic query on city, bedrooms, and rent runs before any model sees a listing.
All three used the full listing prompt and the strong model. The first approach cost an estimated $1.88 across 2,500 calls. The city-filtered approach cost roughly $0.61 with 796 calls. The database query starting point cost about $0.20 with 264 calls. None of them found all 101 matches perfectly — each missed the same one tricky listing.
I measure every subsequent saving against the $0.20 database query starting point, since that is what a careful engineer would build first.
## Letting Code Settle What It Can
The pets field in the dataset is a game-changer. When it is filled in, it explicitly says whether dogs are allowed or not. No model needs to read the listing text for those cases.
Of the 264 pairs that survived the database query, 107 were already decided by the pets field. Seven were rejected because the field listed only cats, and 100 were accepted because dogs were listed. That left only 157 pairs — covering 128 distinct listings — for a model to actually read.
By routing around those 107 pairs entirely, the estimated cost dropped from $0.20 to about $0.11, a 44 percent reduction. Model time fell from 352 seconds to 203 seconds. Precision and recall stayed essentially unchanged because the field check only changed which pairs reached the model; it did not alter the answer for the remaining ones.
This step alone accounted for nearly half of the total cost reduction, and it required zero model calls.
## Does Less Text Mean a Better Answer?
After the pets field settles most decisions, the remaining 157 pairs are ones where the pets field is blank and the agent must read the listing text. For those, I tested a shorter prompt: instead of sending the full listing with title, body, amenities, rent, square footage, city, and pets field, I sent only the title and the body text.
Input tokens dropped by about 31 percent, and the estimated cost fell from $0.114 to $0.098. But the savings were smaller than the token reduction suggests, because output tokens increased from 4,589 to 5,039 — and output tokens cost five times more than input tokens.
The more surprising finding was accuracy. The full listing prompt missed the one true match (a listing that calls itself “a pet friendly community”) every single time across 10 repeated attempts. The shorter prompt — title and body only — found that same match 10 out of 10 times. I have a hypothesis about why: the full prompt included a line stating the pets field was not listed, placed right next to the instruction that unstated pets mean no confirmation. That adjacent cue may have confused the model. I did not test that theory directly, but the effect was clear.
On a second difficult listing — one that mentions “a pet bar for your favorite 4 legged friends” — both prompts split their answers. Neither could reliably tell whether that hints at pet-friendliness or not, and both models answered correctly roughly 7 to 8 times out of 10. That listing is, by nature, a judgment call.
## Reusing Answers Safely
Many of the 157 pairs involve the same listing checked against different renter requirements. A listing that appears in Austin under two different bedroom counts will be asked the same pet question twice.
I added a cache that stores pet decisions keyed by model, listing title, body, pets field, and prompt version. Editing a listing, changing the prompt, or switching models changes the key, so stale answers are never served. The renter requirement is not part of the key because the pet question is the same regardless of which renter is asking.
This cache answered 29 of the 157 pairs from memory. Model calls fell from 157 to 128, and estimated cost dropped from $0.098 to $0.082 — an 8 percent saving with no accuracy loss.
## Using a Cheaper Model with a Strong Backup
The cheapest configuration I tested replaced the strong model entirely with the cheaper model for the pet-reading task. Luna handled 128 listings at a fraction of the cost. In one captured run, Luna made one wrong match; in a rerun, it made none. The estimated cost plummeted to roughly $0.005.
To add a safety net, I set up an escalation rule: Luna answers first, and if its confidence falls below a threshold, the strong model takes a second look. In practice, Luna scored a confidence of 5 on 127 of its 128 questions, and Sol scored 5 on all 157 of its questions. Confidence scores barely varied, so the escalation threshold had almost nothing to trigger on.
In the captured run, the final configuration escalated exactly one answer. Sol re-read the “pet bar” listing, said no with a confidence of 4, and the run had zero wrong matches. Total cost was about $0.008 — roughly 25 times cheaper than the database query starting point and nearly 237 times cheaper than the naive approach of sending every pair to the strong model.
One caveat: a rerun showed Luna answering that same hard listing correctly on its own with a confidence of 5, meaning the escalation did not demonstrably fix anything in that particular run. The result is illustrative, not a validated guarantee.
## Where the Money Actually Went
Starting from the $0.20 database query baseline, the total drop to $0.008 broke down roughly like this: letting code settle the pets field accounted for about 44 percent of the savings. Using the cheaper model with a strong backup accounted for about 39 percent. The shorter prompt contributed roughly 9 percent, and reused answers contributed about 8 percent.
Per 1,000 pairs, estimated cost dropped from $0.080 to $0.0032. Model time fell from 352 seconds to 139 seconds.
The most important finding was not the final number — it was that most model calls were avoidable. Before a model ever reads a listing, code and structured data can settle the vast majority of pairs. The model only needs to handle the cases where structured fields leave ambiguity.
## What to Monitor After Launch
Once the agent is running in production, four metrics are worth tracking continuously:
– **Model calls per pair.** The naive approach produces 1.0 per pair. The query starting point produces about 0.106. The final configuration produces about 0.052. A sudden jump means more pairs are reaching the model than they should.
– **The share of pairs settled by field checks.** This was 93.7 percent in the experiment. A drop suggests the feed format or the renter requirements have changed.
– **Tokens per call, including reasoning tokens.** Prompt changes or model changes show up here first.
– **Precision and recall on a fixed set of pairs with known answers.** Re-running these 2,500 pairs with the final configuration costs less than a cent at the rates quoted above.
Two setup details matter. Register rates with a cost-tracking method so that your tracing tool shows real numbers instead of zeros. Keep the prompt version string in the cache key so that a prompt edit cannot accidentally reuse an old answer.
The confidence scores deserve attention too. If nearly every answer scores the maximum, an escalation rule that triggers below that maximum will rarely fire, leaving the cheaper model to carry every decision with no backup.
## Limitations of This Test
Only one of the 101 true matches depended on reading the listing text. The other 100 came from the structured pets field. This test therefore measures cost reduction far better than it measures the model’s understanding of rental advertisements. The sample is from 2019, so nothing here describes today’s rents or listing styles. The mix of cities was deliberately varied to give the filter something to reject, which also flatters the filter’s performance. A feed focused on a single city would behave differently.
The confidence threshold was not validated in a rigorous tuning exercise. I split the listings into development and test halves, but the development half contained no errors from any configuration, leaving nothing to tune on. The threshold of 5 was chosen after observing which answer carried a confidence of 4. Treat this as an illustration of the approach, not a production-ready setting.
—
## FAQ
**Why did the strong model miss the “pet friendly community” listing while the shorter prompt found it?**
The most likely explanation is that the full prompt included a note saying the pets field was not listed, placed directly next to the instruction that unstated pets should be treated as not confirmed. This adjacent cue may have caused the model to default to “no” regardless of what the title and body said. I did not isolate this variable experimentally, so I cannot confirm it with certainty.
**How much does token count actually vary between runs?**
Token counts are not perfectly deterministic, even with the same prompt and model. In my reruns, the same configuration produced slightly different numbers. One run of the cheaper-model-only configuration made zero wrong matches, while another run made one. Answers can differ run to run when the model does not accept a temperature setting.
**Is caching always safe for this kind of agent?**
Caching is safe when the cache key includes every piece of input that could change the answer. In this experiment, the key included the listing title, body, pets field, and prompt version. The renter requirement was intentionally left out because the pet question does not depend on which renter is asking. A production cache also needs an expiry mechanism and a way to invalidate entries when a listing is edited or removed.
**Why did the Weave usage table show zero tokens and zero cost for every call?**
Weave’s built-in pricing integration did not apply to the models I used, and the experiment did not call the method for adding custom costs. The real token counts were stored in the raw call records, but the usage summary displayed zero. I calculated all cost figures from the stored token counts and the provider’s rate card.
**Would these savings hold for a different domain or a different dataset?**
The savings depend heavily on how much structured data is available before the model sees anything. An agent that processes listings with rich structured fields will see similar savings. An agent where most decisions require reading long, unstructured text will see smaller gains from pre-filtering. The relative ranking of the cheaper model versus the strong model also depends on the specific workload and how often the cheaper model gets it wrong.
**Is it worth keeping a model in the loop at all if a keyword search works?**
For this specific workload, a keyword search for phrases like “pet friendly” and “no pets” agreed with my ground truth on all 128 blank-pets listings. The model is slower and more expensive than a keyword search, and it introduces the risk of wrong answers. The model is justified when the rules are too nuanced for simple keywords — for example, when context, negation, or implied meaning matters. In cases like that, the cost savings documented here still apply.
—
## Conclusion
Reducing LLM costs is not about using a weaker model or asking it to do less. It is about making sure the model only sees the work that actually requires its judgment. In this experiment, structured field checks settled nearly 94 percent of all pairs before a model was ever called, and the remaining work was handed to a cheaper model with a strong backup. The result was a roughly 25-fold cost reduction against a reasonable starting point and a roughly 237-fold reduction against the naive approach — all without losing a single correct answer.
The same principles apply to any agent that processes structured records against a set of rules. Trace your calls, identify what code can settle, route around the model wherever possible, and use the cheapest model that handles the residual ambiguity well enough. Measure everything against a fixed set of labeled examples so you know when quality slips.
Cost optimization is not a one-time exercise. Traffic patterns change, prompts get revised, new models arrive, and the balance between pre-filtering and model inference shifts. The four metrics outlined in this article — calls per pair, field-check settlement rate, tokens per call, and precision and recall on a fixed sample — form a lightweight monitoring framework that catches drift early.
The most surprising result was not the magnitude of the savings. It was how little of the original work actually required a model at all. The hardest part of cutting cost is not building a fancier model — it is building the discipline to let code do what code does best, and reserving the model for the questions that genuinely need it.
Thank you for reading.



