Skip to content

Asana Cuts Browser Agent Costs by Optimizing History and Caching

Asana’s experiments show that browsing history management can determine costs and task completion. The 76-fold reduction combines workflow optimization with a model change, rather than caching alone.

•7 min read
Share:
Asana Cuts Browser Agent Costs by Optimizing History and Caching

Asana has deployed changes to the browser agent in StackAI following a study involving 144 runs, according to a case study published by OpenAI on October 9, 2026. The optimized configuration using GPT-6.1 Sol achieved an average estimated model cost of $0.47 and a runtime of about four minutes per run, reportedly making it 76 times cheaper and five times faster than the original production configuration, which used a different model.

The notable point is how history is managed: retaining more information and modifying screenshots less frequently helps the agent take advantage of caching while avoiding revisits to pages it has already read. However, the 76-fold improvement combines workflow changes with model selection, rather than reflecting the effect of caching alone.

Browsing History Becomes a Cost Bottleneck

StackAI, the platform acquired by Asana, lets users build workflows that navigate websites, fill out forms and collect information without writing code. With this type of agent, each step can add page text and screenshots to the context sent to the model.

According to OpenAI, Frank Hidalgo, CTO of StackAI at Asana, used GPT-6 Astra in Codex to examine the source code and how the agent constructed each request. The system already cached fixed instructions and tool definitions, but not the growing browsing history. As a result, the history was resent at uncached input prices.

Another issue meant that simply adding caching was not enough: the agent deleted old images and trimmed text at almost every step. These operations continually changed the history. They could also remove necessary facts, forcing the agent to revisit pages it had already read.

In terms of how it works, input caching is beneficial when repeated portions of content are reused. Keeping history stable for longer makes that possible; continually trimming and editing history can reduce the benefits even if the overall context is shorter.

Changes Tested Across 144 Runs

Comparison of deleting images at every step versus pruning images in batches

Hidalgo chose three approaches to test: extending caching to browsing history, increasing the amount of text retained and deleting screenshots in batches rather than at every step.

GPT-6 Astra then refactored the code so that the same frontend and backend could support multiple workflows with separate configurations running in parallel. The study compared two history budgets, 120,000 and 480,000 characters, along with six caching and screenshot policies. Each combination was run three times on each of four models, yielding 144 runs.

All configurations performed the same task: collecting six fields of information for each of 32 books from a public demo catalog. This was a specific scope, not a test suite covering every website or business process.

The selected policy allowed screenshots to accumulate to 20 before retaining only the most recent one. Combined with the larger text budget, this kept the earlier history unchanged across multiple consecutive steps.

Requests, data traces and results were recorded in Command, Asana’s software delivery platform. According to the source, the results were turned into tickets and pull requests, then deployed to production. This approach to recording traces also relates to the need to reproduce failures in product testing with AI agents: the final output alone is not enough to explain a failed run.

The 76-Fold Reduction Must Be Distinguished from Same-Model Gains

The most striking figure compares the original configuration on Model B with the optimized configuration on GPT-6.1 Sol. Model B is anonymized in the publication, so there is no basis for identifying a specific provider or product.

For Model B alone, estimated model costs fell from at least $36.21 to $1.24 per run, a roughly 29-fold reduction. The baseline is a lower bound because some runs reached the step limit before completing.

On GPT-6.1 Sol, with the larger history budget held constant, the new caching and screenshot policy reduced costs from $1.97 to $0.47, roughly a fourfold reduction. The source states that 89% of input came from the cache, at a unit price equal to 5% of uncached input in this measurement.

The 76-fold reduction therefore should not be interpreted as meaning that simply enabling caching will produce similar results. It reflects a combination of changes and uses a baseline configuration that included incomplete runs. The reported amounts are also estimated model costs, not the full cost of browser operation, infrastructure and staffing.

Retaining Context Also Affects Task Completion

Retaining enough context helps avoid rereading information already collected

Cost was not the only outcome. With GPT-6.1 Sol, increasing the history budget from the smaller to the larger setting raised the number of runs that produced an answer from three out of 18 to all 18; every answer in the larger-budget group was correct for the test task.

The results show that reducing context does not automatically save money. If an agent loses information and has to reread a page, or fails to complete the work, a lower cost per request does not necessarily translate into a lower cost per successful task.

This is also why model capabilities need to be distinguished from the runtime environment, as discussed in the article on evaluating GPT-6 Astra before deployment. In the Asana study, Astra helped with investigation and experimentation; Sol was the model running the browser configuration that achieved the reported results.

Deployed, but Generalizability Remains Limited

Asana says it has released the browser navigation changes in StackAI and is developing tools to make the experiments easier to repeat. The company plans to incorporate comparisons of cost, runtime and answer quality into the platform’s evaluation system.

The results are currently presented in a case study published by OpenAI, not an independent evaluation. Each configuration had three repeat runs, all performing the same data collection task. They have not established savings for websites requiring login, changing interfaces or workflows involving data-writing actions.

Moving agents from experimentation into operation therefore remains tied to the challenge of governing internal AI applications. This study provides concrete evidence about context and caching optimization, but does not demonstrate access control capabilities or safety for every browser task.

Frequently Asked Questions

Is a budget of 480,000 characters equivalent to 480,000 tokens?

No. Characters and tokens are different units; the token count depends on the content and the model’s tokenizer. The study describes the history budget in characters, so it cannot be directly converted into a token limit or context window.

Can the policy of retaining 20 screenshots be applied immediately to every agent?

This was the configuration selected for Asana’s test task, not a universally optimal threshold. Tasks that rely heavily on visual information or involve different numbers of steps need separate evaluations of quality, time and cost.

Further Reading

Share: