# Retrieval, Action, and the Space Between: A Practical Experiment in Python
## TL;DR
– A pure Python implementation was built to compare retrieval-only, action-only, and hybrid systems.
– Nine tasks were run across all three systems, with expected ticket states frozen before execution.
– Two real bugs were discovered during the build, and both are documented with terminal output showing the failure and the fix.
– The experiment demonstrates that retrieval alone cannot perform actions, actions alone cannot answer knowledge questions, and only the combined system handles tasks requiring both.
– No external language model, vector database, or embedding API was used—only the Python standard library.
—
## Why This Experiment Exists
A common pattern in modern AI development is to assume that combining retrieval with action automatically produces a better system. The idea is appealing: find the right information, then do something with it. But the assumption rarely comes with proof.
This experiment was designed to test that assumption directly. Instead of arguing about definitions or comparing conceptual architectures, the goal was to build working systems, feed them identical tasks, and measure what actually happens. The question at the center was simple: does connecting retrieval to action change the outcome, and if so, under what conditions?
—
## Three Systems, One Test Suite
Three separate implementations were built. Each was given the same nine tasks, and each was evaluated against a known expected final state.
**System 1: Retrieval-Only**
This system can search a document corpus and return relevant passages. It cannot modify any external state. If a task asks it to update a record or change a value, it explicitly reports that it lacks the ability to do so.
**System 2: Action-Only Planner**
This system can read task instructions, parse them, and execute actions against a ticket store. It can change priorities, update statuses, assign owners, and set categories. However, it has no access to any documentation or knowledge base. If a task requires domain-specific information that isn’t stated directly in the prompt, it cannot find that information.
**System 3: Hybrid System**
This system connects the two. It retrieves relevant documents from the corpus and passes them to the action planner, which then uses that retrieved context to decide which action to take. This is the only system that has access to both information and the ability to act on it.
All three systems were implemented in standard Python 3.12 with no external dependencies. No language model was involved. No vector database or embedding service was used. This was a deliberate choice—keeping the implementation small and transparent makes it possible to trace exactly why a system succeeds or fails at each step.
—
## Architecture Overview
The project was organized into five distinct modules, each with a single responsibility.
The **retriever module** contains the document corpus and the search logic. The corpus consists of 334 chunks drawn from technical articles covering retrieval-augmented generation, AI agents, data science, and Python programming, totaling roughly 59,000 words. The retrieval mechanism uses term frequency–inverse document frequency with cosine similarity to rank chunks. There is no stemming, no lemmatization, and no approximate nearest-neighbor index—just straightforward token matching and scoring. A single query takes roughly 1.4 milliseconds once the index is built.
The **environment module** holds the ticket store. Four tickets are tracked, and four actions can be performed on them: changing priority, changing status, assigning an owner, and setting a category. Every action is logged, and the final state of each ticket is checked independently of what the system reports.
The **agent planner module** handles action logic. It uses regular expressions to extract a ticket identifier and field values from the task description, then applies the corresponding action to the ticket store. It does not import the retriever and has no awareness of the document corpus.
The **retrieval-only system module** imports the retriever but has no access to the ticket environment. It answers knowledge questions and refuses action requests.
The **hybrid system module** is the only file that imports both the retriever and the agent planner. It calls retrieval first, then passes the retrieved chunks to the planner as context, and then checks the ticket state after the action is completed.
This strict separation matters. When the hybrid system fails or succeeds, it is possible to see exactly where the retrieved information was used to drive an action decision—and where it was not.
—
## The Nine Tasks
The tasks were divided into three groups of three. The expected final ticket states were defined before any code was executed, so there was no ambiguity about what a correct outcome looks like.
**Knowledge-Only Tasks (A1–A3):** These ask questions that can be answered entirely from the document corpus. No ticket modifications are required. An example asks about the main chunking strategies discussed in the material.
**Action-Only Tasks (B1–B3):** These instruct the system to make specific changes to a ticket. An example assigns a ticket to a person and sets its priority. No document lookups are needed because all the information is in the task prompt itself.
**Knowledge + Action Tasks (C1–C3):** These require both. The system must first search the corpus to determine the correct category for a ticket, then update the ticket with that category. The ticket description alone is insufficient—retrieved documentation is needed to make the right decision.
Each task starts from a fresh copy of the ticket store. No state carries over between runs, which ensures that earlier tasks do not influence later results.
—
## What Happened When It Was Run: The First Bug
The initial hybrid planner had a straightforward flaw. It looked for a ticket ID before doing anything else. If no ticket ID was found, it stopped immediately—regardless of whether retrieval had already returned useful information.
This was fine for action tasks, where a ticket ID is expected. But for knowledge-only tasks, which contain no ticket ID at all, the hybrid system failed before it ever used the retrieved documents. The plain retrieval system handled these tasks correctly, but the hybrid system, which in theory had more capabilities, failed on tasks that the simpler system could handle.
The terminal output showed the failure clearly: the hybrid system reported that no ticket ID was found in the task, even though it had retrieved the correct chunks from the corpus.
The fix was to add a check that determines whether an action is actually required before running the workflow. If the task does not contain both an action verb and a ticket identifier, the system uses the retrieved context to answer directly rather than attempting to modify a ticket. After the fix, the hybrid system passed all knowledge-only tasks.
The broader lesson is that combining two working components does not automatically produce a working combined system. The interface between them matters, and assumptions baked into one component can break the other when they are connected.
—
## The Second Bug
After fixing the first issue, two action-only tasks failed. Both the standalone agent and the hybrid system reported success, but the actual ticket state did not match what was expected.
In one case, the ticket was assigned the correct priority but the assignee field was empty. In the other, a status change was applied but the wording of the task did not match the parser’s expectations.
The root cause was in the regular expression patterns used to parse task instructions. The assignment regex expected the form “assign to Alice,” but the task used “assign T101 to Alice”—with a ticket ID inserted between the verb and the preposition. The parser skipped the instruction entirely and reported the actions it did manage to run as a success, without checking whether it had processed every part of the task.
A ground-truth check against the actual ticket state caught this discrepancy. The system’s own output said “pass,” but the ticket said the assignee was still unset. This mismatch between self-reported results and actual state is an important reason to validate outcomes independently rather than trusting the system’s summary.
Both regular expressions were updated to handle the task wording. After the changes, the parser was run against all nine tasks again from a clean environment to confirm the fix worked consistently.
—
## Final Results
After both bugs were resolved and the environment was reset, all 27 executions were rerun and validated against the expected ticket states.
The retrieval-only system passed the three knowledge tasks but failed all six action tasks—expected, since it has no code to modify tickets. The action-only planner passed the three action tasks but failed the three knowledge tasks—expected, since it has no access to the corpus. The hybrid system passed all nine tasks. It was the only system that could both find the needed information and apply it to make a correct ticket update.
The results confirm that the value of combining retrieval with action depends entirely on what the task requires. For knowledge-only questions, retrieval alone is sufficient. For action-only instructions, an action-capable system alone is sufficient. For tasks that require both information and a subsequent action, only the connected system works.
—
## Surprising Findings: Noisy Retrieval and Majority Voting
In the knowledge-plus-action tasks, the top retrieved result was sometimes incorrect. This happened because the task instructions themselves contained words like “knowledge base” and “appropriate category,” which also appear frequently in the corpus. The retriever treated the entire task text as a query and matched those instructional phrases, causing a relevant but misaligned document to rank first.
However, the category resolver uses all three retrieved chunks, not just the top one. It counts which topic group appears most frequently across the results and selects the winner. In the cases where the first result was wrong, the second and third results carried the correct signal, and the majority vote still produced the right category.
This reveals an interesting property of the system: the retrieval quality of individual queries matters less than the aggregation strategy. A noisy first result does not necessarily corrupt the final answer if the resolver looks at multiple candidates. That said, this was observed across only three tasks with one corpus, so it would be unwise to generalize without further testing.
—
## Performance Characteristics
The main operations were timed on a single machine running Python 3.12 with no GPU. The retriever initialization—loading all chunks and building the TF-IDF index—took roughly 29 milliseconds. Once the index was built, each retrieval query took about 1.4 milliseconds. Environment actions completed in approximately 0.001 milliseconds each. The agent planner, which parses the task text with regexes and applies one or two actions, took about 0.009 milliseconds.
A full run of the retrieval-only system took roughly 1.3 milliseconds. A full run of the hybrid system with a prebuilt index took roughly 1.5 milliseconds. The difference between the two is almost entirely the category resolution and the action execution, both of which are very fast.
One observation worth noting: in this experiment, a new system instance was created for every task to ensure a clean ticket state each time. That means the TF-IDF index was rebuilt multiple times during testing. In a production setting, the index would be built once and reused, making retrieval costs negligible across many queries.
—
## Limitations
Several limitations should be acknowledged. The retrieval method uses TF-IDF, which relies on exact word matching and has no understanding of semantic similarity. Queries that use synonyms or different phrasing for the same concept may not match relevant documents. The noisy top-1 results in some tasks were a direct consequence of this.
The action planner is constrained by the regular expressions that were written. Sentence structures that differ from the patterns in the test set will not be parsed correctly. The parser was tested against a small number of phrasings, and its robustness to varied wording is unknown.
The category resolver uses a simple majority vote with no confidence threshold. A 2-to-1 split and a 3-to-0 split both produce the same output, even though the certainty is clearly different. In a higher-stakes system, it would be important to flag uncertain cases rather than act on thin evidence.
Finally, the test set is small—nine tasks from one corpus. This was sufficient to demonstrate the core separation between retrieval and action and to surface two implementation bugs. It is not sufficient to support broad claims about how these architectures would perform at scale or with different domains.
—
## Frequently Asked Questions
**What is the difference between retrieval and action in this context?**
Retrieval is the process of finding relevant information in a document corpus. Action is the process of modifying an external state, such as updating a ticket field. They are complementary capabilities that solve different parts of a task.
**Why was no language model used?**
A language model was intentionally excluded so that every part of the pipeline—retrieval, decision-making, and action—could be written and inspected directly. This makes debugging much simpler, since failures can be traced to specific lines of code rather than to opaque model behavior.
**What retrieval method was used?**
Term frequency–inverse document frequency with cosine similarity. Documents are tokenized, stopwords are removed, and query vectors are compared against document vectors. No embeddings or neural networks are involved.
**How were the tasks evaluated?**
Each task was checked against a pre-defined expected ticket state. The system’s own pass/fail report was not trusted as the final verdict. Instead, the actual ticket fields were inspected independently after each run.
**Why did the hybrid system sometimes retrieve the wrong document first?**
Task instructions often contain words like “knowledge base” and “category,” which overlap with terms common in the corpus. The retriever sees the full task text as a query and matches these terms, which can push instructional text to the top of the results.
**Does majority voting over retrieved chunks always fix noisy retrieval?**
Not necessarily. In this experiment, two out of three retrieved chunks pointed to the correct category, so majority voting worked. If retrieval quality degrades further—for example, if most chunks are irrelevant—majority voting alone will not save the result.
**Can the results be reproduced?**
Yes. The implementation uses only the Python standard library and a local corpus. No API keys, network calls, or external services are required once the corpus is prepared.
**What are the most important takeaways?**
First, retrieval and action are distinct capabilities that solve different problems. Second, connecting them is not automatic—bugs can appear at the boundary. Third, the value of a hybrid system depends on the task: it is essential when both information and action are needed, but unnecessary when only one is required.
—
## Conclusion
This experiment set out to test a practical question: what happens when retrieval and action are kept separate versus connected? The answer turned out to be nuanced rather than dramatic. Each system excelled at the tasks that matched its capabilities and failed at the tasks that exceeded them. The hybrid system was the only one that could handle knowledge-plus-action tasks, but only after two bugs in the connecting logic were fixed.
The most striking finding was not about performance or accuracy. It was about where failures occur. The retrieval code worked. The action code worked. But when the two were wired together—when the output of one became the input of the other—new failure modes appeared. The planner assumed every task was an action request. The parser did not handle the specific wording used in the task descriptions. These were not deep architectural failures. They were small, local bugs that only surfaced at the boundary between components.
That boundary is where the real engineering lives. Retrieval and action each look simple in isolation. Connecting them reveals the seams. The experiment shows that building a working hybrid system requires attention not just to the individual parts, but to how they talk to each other—what gets passed, how it is parsed, and whether the final state actually matches what was intended.
The results also reinforce a practical principle: measure outcomes against ground truth, not against the system’s own self-report. One of the bugs found here would have been invisible if only the planner’s success message had been trusted. Independent validation of the actual state is essential.
This kind of hands-on building surfaces truths that conceptual comparisons cannot. The code is small enough to read in full, the tasks are few enough to trace by hand, and the results are specific enough to be meaningful. The separation between finding information and doing something with it is real, and it matters—but only when the task demands both.
Thank you for reading



