Requests served from cache
Every request in the published replay preview was a cache hit.
Your agent already figured it out.
Let it remember. Cache successful model responses and replay repeated workflows through an OpenAI-compatible API.
$0 upstream model cost on exact cache hits. First runs, cache misses, and the infrastructure that executes your tasks still have their own costs.
Repeated requests. Reused intelligence.
The benchmark preview published in our GitHub repository shows what happens when the work is already cached.
Every request in the published replay preview was a cache hit.
Exact cached responses avoid another upstream model call.
The runtime improvement reported in the repository preview.
Source: the repository’s published benchmark preview. These are reported replay results, not a guarantee for every workload. Savings depend on cache hits; model-response replay does not remove browser, sandbox, network, or execution costs.
Inspect the benchmark ↗A small cache between your agent and its model provider. One familiar API, with successful responses ready for the next run.
Point your OpenAI-compatible client at the proxy. Keep your models and tools, with OpenRouter or your own compatible upstream provider.
Successful responses are stored. Repeat an identical request and the cache returns the response locally, without asking the upstream model again.
Enable optional JEV to judge whether a previous response can serve a new request. If it cannot, the request continues to your generation provider.
Explore JEV configuration ↗Choose the Node file cache or Python SQLite server. Configure expiration and model filters, inspect hit/miss headers, and bypass caching per request.
Read the configuration reference ↗A great agent shouldn’t have to rediscover the same answer every time. Do the work once. Give the next run a head start.
Start with the workflows in the repository. Reuse a plan, script, or response while your tools handle execution.
Reuse browser plans and tool-call responses for repeated Browserbase workflows.
Browserbase example ↗Cache a website-generation response, then reuse it when the model request is unchanged.
Website example ↗Reuse a download script or command plan and let your execution environment run it.
Download example ↗Run the open-source cache locally, connect your provider, and point your client at its OpenAI-compatible endpoint.
Use your own upstream API key. The cache works in front of OpenRouter or an OpenAI-compatible provider.
Run the Node proxy, then use the local base URL in your agent or SDK.
Make a non-streaming request to populate the cache. Send the same request again and inspect the response header.
export UPSTREAM_BASE_URL=https://openrouter.ai/api/v1
export UPSTREAM_API_KEY="your-provider-key"
npx -y github:rohanarun/computer-use-cache start
# In your agent's shell:
export OPENAI_BASE_URL=http://127.0.0.1:8000/v1
export OPENAI_API_KEY="$UPSTREAM_API_KEY"import OpenAI from "openai";
const client = new OpenAI({
baseURL: "http://127.0.0.1:8000/v1",
apiKey: process.env.UPSTREAM_API_KEY,
});
// Keep your existing chat.completions calls.
// Repeat an identical request to get a cache hit.# Generate setup files for your agent:
npx -y github:rohanarun/computer-use-cache install codex
# Or: claude-code, cursor, openclaw, hermes, all
# Inspect savings and setup:
npx -y github:rohanarun/computer-use-cache stats
npx -y github:rohanarun/computer-use-cache doctorExact cache hits stay local. Optional JEV reuse makes a separate model judgment and can incur provider costs. For fresh observations or actions that must run again, use cache: false.
Read the code. Run the examples. Inspect the benchmark preview. Computer-Use Cache is MIT licensed and built to fit into your own agent infrastructure.
Explore the repository ↗An OpenAI-compatible cache for repeated computer-use and agent workflows.
An exact cache hit reuses a stored model response and avoids another upstream model call. Initial generation and cache misses still use your provider. Hosting, browsers, sandboxes, network traffic, and task execution can still cost money.
The cache returns a model response, such as a plan or tool call. Your agent and its tools still execute the task. A cache hit alone is not evidence that an external action completed.
By default, it is an exact cache miss and goes to your upstream model. With JEV enabled, a model first judges whether an eligible cached response can be reused unchanged. If reuse is rejected or the judgment fails, the request falls through to normal generation.
Exact cache hits can stream cached text through Server-Sent Events. Streaming misses pass through upstream and are not stored, so make the first request non-streaming to populate the cache.
Bring Computer-Use Cache to your agent.