On this article, you’ll be taught what software calling and code execution are as agent motion primitives, how they differ mechanically, and when to decide on one over the opposite.
Subjects we’ll cowl embody:
- How software calling works below the hood, and why it stays the fitting alternative for single, time-sensitive lookups.
- How code execution by way of Programmatic Device Calling differs from commonplace software calling, and what measurable advantages it provides for fan-out and aggregation duties.
- A sensible choice framework for selecting between the 2 primitives primarily based on name rely, information sensitivity, latency, infrastructure, and auditability wants.
Image an agent asking one simple-sounding query: which of twenty staff went over their Q3 journey finances. To reply it, the agent wants every individual’s expense line objects, each flight, lodge, and meal receipt, in contrast towards a finances restrict tied to their stage. Constructed the apparent manner, with the mannequin calling a software for every individual’s bills one by one, that’s twenty separate software calls, every returning fifty to 100 line objects, and each single a kind of objects has to move via the mannequin’s context simply so it may be added up. That’s over 2,000 line objects and greater than 50KB of uncooked information the mannequin by no means truly wanted to learn — it wanted a sum.
That’s the actual value hiding behind a design choice most agent tutorials skip previous totally: how does an agent truly take motion on the earth. There are two actual solutions, software calling and code execution, and which one you attain for isn’t a method choice — it’s an architectural alternative with measurable penalties for value, latency, and accuracy. This text breaks down each motion primitives for AI brokers intimately, builds an actual, runnable instance of every utilizing the identical underlying software, and closes with an sincere, numbers-backed framework for selecting between them. Should you haven’t constructed a fundamental tool-calling agent but, try this text, Easy Agentic Tool Calling with Gemma 4 — it’s the pure place to start out earlier than this one.
What Is an Motion Primitive, and Why Does the Selection Matter?
An motion primitive is the elemental mechanism by which a language mannequin turns a call into an actual impact on the earth — a database write, an API name, a file learn. Each agent framework, no matter else it does, is constructed on high of certainly one of these primitives at its core.
Device calling is the primitive most individuals be taught first: the mannequin produces one structured request at a time, a number software executes it, and the consequence comes again into the dialog earlier than the mannequin decides what to do subsequent. Code execution is the newer different: as an alternative of requesting one motion and ready, the mannequin writes an precise program — in Python or TypeScript — that performs a number of actions in sequence or in parallel, and solely this system’s ultimate output returns to the mannequin.
Neither one is a wrapper across the different, and neither has quietly changed the opposite. They’re genuinely totally different mechanisms with totally different failure modes, totally different infrastructure necessities, and totally different value profiles, and the remainder of this text is about understanding each properly sufficient to choose accurately.
Device Calling
It’s value understanding what’s truly taking place beneath a software name, as a result of the mechanics clarify each its strengths and its actual limitations. In accordance with Cloudflare’s detailed breakdown of the method, a mannequin producing a software name doesn’t produce unusual textual content. It’s been particularly skilled to output a pair of particular tokens — one signaling “the next is a software name” and one other marking its finish — with a JSON payload describing the software title and arguments sitting between them. The appliance operating the mannequin watches for these tokens, pauses technology the second it sees the closing one, parses the JSON towards a schema you outlined, truly executes the decision, and feeds the consequence again into the dialog as if it had been the following factor the person stated.
That’s a clear, auditable, one-step-at-a-time loop, and it’s precisely why software calling turned the default. Each motion is a discrete, loggable occasion. Each result’s one thing the mannequin instantly sees and might purpose about in pure language earlier than deciding what occurs subsequent.
Code Execution
Code execution takes a unique beginning place totally: as an alternative of asking the mannequin to explain an motion in a constrained JSON format, you let it write precise code that performs the motion, operating in a sandboxed setting separate from the mannequin itself. Anthropic’s original code-execution-with-MCP pattern frames this exactly as presenting your instruments as a code API moderately than a set of instantly callable features, so the mannequin can write a script that imports precisely the instruments it wants and calls them the best way it might name every other perform.
The mechanism that makes this genuinely totally different — not only a relabeled software name — is what Anthropic now calls Programmatic Tool Calling, launched alongside two companion options in November 2025. Relatively than every software consequence flowing again via the mannequin one by one, you mark particular instruments as callable from code by including an allowed_callers area to their definition, and add a code_execution software to the request. When the mannequin desires to behave, it writes a full script — loops, conditionals, error dealing with, and all — that calls these instruments instantly inside a sandboxed execution setting. Every particular person software name the script makes nonetheless executes precisely the best way it might in unusual software calling; you continue to obtain a request and return a consequence, however that result’s intercepted and processed by the operating script moderately than being pushed into the mannequin’s context. Solely when the script finishes does its ultimate output — and nothing else — return to the mannequin.
That’s the complete distinction in a single sentence: software calling places each intermediate lead to entrance of the mannequin; code execution lets the mannequin resolve, via the code it writes, precisely what makes it again.
A side-by-side stream diagram of Device Calling and Code Execution (click on to enlarge)
Device Calling for a Single, Time-Delicate Lookup
Idea is simpler to belief as soon as it’s operating towards an actual API, so each examples on this article use the identical software — a get_weather perform backed by Open-Meteo, a free climate API that wants no API key in any respect, solely an Anthropic API key to run the agent itself.
Begin with the case software calling is clearly proper for: a single query that wants one lookup and a natural-language reply — “what’s the climate like in London proper now.”
|
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 |
import json import requests from anthropic import Anthropic
consumer = Anthropic() # reads ANTHROPIC_API_KEY from the setting
def get_weather(metropolis: str) -> dict: “”“Search for a metropolis’s coordinates, then fetch its present temperature and this week’s each day highs from Open-Meteo’s free, keyless API.”“” geo = requests.get( “https://geocoding-api.open-meteo.com/v1/search”, params={“title”: metropolis, “rely”: 1}, ).json() if not geo.get(“outcomes”): return {“error”: f“Couldn’t discover a location named ‘{metropolis}'”} lat = geo[“results”][0][“latitude”] lon = geo[“results”][0][“longitude”]
forecast = requests.get( “https://api.open-meteo.com/v1/forecast”, params={ “latitude”: lat, “longitude”: lon, “present”: “temperature_2m”, “each day”: “temperature_2m_max”, “timezone”: “auto”, }, ).json()
return { “metropolis”: metropolis, “current_temp_c”: forecast[“current”][“temperature_2m”], “week_high_temps_c”: forecast[“daily”][“temperature_2m_max”], “unit”: “celsius”, }
weather_tool = { “title”: “get_weather”, “description”: ( “Get the present temperature and this week’s each day excessive “ “temperatures for a metropolis. Returns JSON with metropolis, “ “current_temp_c, week_high_temps_c (7 each day highs), and unit.” ), “input_schema”: { “kind”: “object”, “properties”: { “metropolis”: {“kind”: “string”, “description”: “Metropolis title, e.g. ‘Lagos'”} }, “required”: [“city”], }, }
messages = [{“role”: “user”, “content”: “What’s the weather like in London right now?”}]
response = consumer.messages.create( mannequin=“claude-sonnet-5”, max_tokens=1024, instruments=[weather_tool], messages=messages, )
# Maintain resolving software calls till Claude produces a ultimate textual content reply whereas response.stop_reason == “tool_use”: messages.append({“position”: “assistant”, “content material”: response.content material}) tool_results = []
for block in response.content material: if block.kind == “tool_use” and block.title == “get_weather”: consequence = get_weather(**block.enter) tool_results.append({ “kind”: “tool_result”, “tool_use_id”: block.id, “content material”: json.dumps(consequence), })
messages.append({“position”: “person”, “content material”: tool_results}) response = consumer.messages.create( mannequin=“claude-sonnet-5”, max_tokens=1024, instruments=[weather_tool], messages=messages, )
for block in response.content material: if block.kind == “textual content”: print(block.textual content) |
Strolling via what issues right here: get_weather itself is unusual Python — nothing agent-specific about it — it geocodes a metropolis title and pulls each the present temperature and the week’s each day highs in a single request. The weather_tool dictionary is the schema Claude truly sees, and the outline issues greater than it appears to be like — a obscure description is among the commonest causes of a mannequin calling a software with the unsuitable arguments. The whereas response.stop_reason == “tool_use” loop is the actual mechanical coronary heart of normal software calling: each time Claude requests the software, your code has to truly run it, wrap the consequence as a tool_result block, append it to the dialog, and name the API once more — and this repeats for as many software calls as the duty wants. For a single lookup like this one, that’s one move via the loop and carried out, which is precisely why software calling matches this case properly: one name, one consequence, and a consequence small and related sufficient that Claude genuinely advantages from seeing it instantly earlier than writing a natural-language reply.
Code Execution for Fan-Out and Aggregation
Now change the query, utilizing the very same get_weather perform — utterly unchanged: “given these fifteen cities, which one could have the coldest excessive temperature this week, and what’s the common weekly excessive throughout all of them?”
Run that via the tool-calling loop above and also you’d get fifteen separate software calls, fifteen full JSON payloads of each day temperatures pushed into Claude’s context, and Claude would then must manually examine and common them in pure language — sluggish, token-expensive, and precisely the sort of arithmetic a mannequin is extra error-prone at than a for-loop is. That is exactly the case Programmatic Device Calling was constructed for.
|
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 |
import json from anthropic import Anthropic
consumer = Anthropic()
# Similar get_weather perform from the earlier instance, unchanged
weather_tool = { “title”: “get_weather”, “description”: ( “Get the present temperature and this week’s each day excessive “ “temperatures for a metropolis. Returns JSON with metropolis, “ “current_temp_c, week_high_temps_c (7 each day highs), and unit.” ), “input_schema”: { “kind”: “object”, “properties”: { “metropolis”: {“kind”: “string”, “description”: “Metropolis title, e.g. ‘Lagos'”} }, “required”: [“city”], }, # That is the one line that adjustments the primitive: it opts the software # into being referred to as from inside generated code, not solely instantly # by the mannequin “allowed_callers”: [“code_execution_20250825”], }
code_execution_tool = {“kind”: “code_execution_20250825”, “title”: “code_execution”}
cities = [ “Lagos”, “Nairobi”, “Cairo”, “Accra”, “Kigali”, “Casablanca”, “Addis Ababa”, “Dakar”, “Tunis”, “Kampala”, “Harare”, “Lusaka”, “Maputo”, “Windhoek”, “Gaborone”, ]
messages = [{ “role”: “user”, “content”: ( f“Given these cities: {‘, ‘.join(cities)}, which one will have “ “the coldest high temperature this week, and what’s the average “ “weekly high across all of them? Use the get_weather tool.” ), }]
response = consumer.beta.messages.create( betas=[“advanced-tool-use-2025-11-20”], mannequin=“claude-sonnet-5”, max_tokens=2048, instruments=[code_execution_tool, weather_tool], messages=messages, )
# The loop appears to be like just like commonplace software calling, however now some # tool_use blocks carry a “caller” area, that means the request got here # from inside Claude’s generated script moderately than from Claude instantly whereas response.stop_reason == “tool_use”: messages.append({“position”: “assistant”, “content material”: response.content material}) tool_results = []
for block in response.content material: if block.kind == “tool_use” and block.title == “get_weather”: consequence = get_weather(**block.enter) tool_results.append({ “kind”: “tool_result”, “tool_use_id”: block.id, “content material”: json.dumps(consequence), })
if tool_results: messages.append({“position”: “person”, “content material”: tool_results})
response = consumer.beta.messages.create( betas=[“advanced-tool-use-2025-11-20”], mannequin=“claude-sonnet-5”, max_tokens=2048, instruments=[code_execution_tool, weather_tool], messages=messages, )
for block in response.content material: if block.kind == “textual content”: print(block.textual content) |
The only most essential line on this entire script is “allowed_callers”: [“code_execution_20250825”]. With out it, the software behaves precisely because it did within the earlier instance — callable solely instantly by the mannequin. With it added, Claude positive factors the choice to write down a script that calls get_weather fifteen instances itself, probably in parallel utilizing asyncio.collect, sum and kind the outcomes, and print solely the ultimate reply — the coldest metropolis and the common — to plain output. Your Python code doesn’t change the way it responds to particular person software calls in any respect; that a part of the loop appears to be like almost similar to the tool-calling instance. What adjustments is invisible out of your aspect of the API: fourteen of these fifteen climate lookups, and each intermediate comparability between them, by no means contact Claude’s context.
Claude solely ever sees the 2 numbers it truly requested for. Since this makes use of a beta function, it’s value double-checking the precise beta header string and block-handling particulars towards Anthropic’s current documentation earlier than counting on it in manufacturing, as beta APIs are the a part of any platform probably to shift.
Why Code Execution Wins at Scale
The climate instance makes the mechanism seen, but it surely’s value backing this up with actual, revealed figures moderately than instinct alone. Anthropic’s authentic code-execution-with-MCP sample took an actual Google Drive-to-Salesforce workflow from 150,000 tokens right down to 2,000 — a 98.7% reduction — just by retaining a full assembly transcript contained in the execution setting as an alternative of routing it via the mannequin twice.
Programmatic Device Calling’s personal inside benchmarking, reported directly by Anthropic, discovered common token utilization on advanced analysis duties dropped from 43,588 to 27,297 — a 37% discount — whereas accuracy on the GAIA benchmark truly improved, rising from 46.5% to 51.2%, and inside information retrieval accuracy rose from 25.6% to twenty-eight.5%. That final element issues greater than the token financial savings alone: this isn’t purely a price optimization. Offloading orchestration logic to precise code moderately than asking a mannequin to trace it via pure language measurably reduces the sort of errors that come from a mannequin shedding observe of a dozen intermediate values it’s attempting to match in its head.
The tutorial consequence beneath all of this predates Anthropic’s personal tooling. The unique CodeAct paper from Wang and colleagues in 2024 discovered that brokers taking motion via executable code, moderately than JSON-formatted software calls, succeeded as much as 20% extra typically on advanced, multi-step duties. Code execution isn’t a current product function bolted onto an current thought — it’s a research-backed sample that the key labs have spent the previous two years turning into manufacturing infrastructure.
The place Device Calling Nonetheless Wins
The numbers above could make code execution appear like an unconditional improve, and it isn’t one. There’s an actual, sincere case for sticking with plain software calling in a significant set of conditions.
Single-call duties are the clearest case. The Lagos climate instance earlier on this article positive factors nothing from a sandbox — one name, one small consequence, and the overhead of spinning up a code execution setting provides latency with out including any actual profit. Duties the place the mannequin genuinely must purpose over an intermediate lead to pure language are the second case: if the precise level of a step is for the mannequin to note one thing refined in a doc or a dataset and reply to it conversationally, filtering that information away in a sandbox defeats the aim. Easier infrastructure is an actual, sensible issue too — a group with out an current safe sandboxing setup takes on actual operational value standing one up, and that value must be weighed towards the financial savings, not assumed away. And auditability issues greater than it will get credit score for: a software name is one clear, loggable occasion with a reputation and a set of arguments, whereas reasoning about precisely what a generated script did internally — particularly after the actual fact, throughout an incident — is a genuinely more durable debugging drawback.
Resolution Framework: Selecting the Proper Primitive
Pulling all the pieces above into one sensible reference:
| Issue | Favors software calling | Favors code execution |
|---|---|---|
| Variety of calls wanted | One, or a small, fastened few | A number of, particularly with fan-out or aggregation |
| What occurs to outcomes | The mannequin must learn and purpose over them instantly | They only should be filtered, summed, or in contrast |
| Information sensitivity | Low — nothing problematic in regards to the mannequin seeing it | Excessive — PII or massive payloads higher stored out of context |
| Latency tolerance | Tight — sandbox startup isn’t value paying for | Workflow already entails a number of round-trips anyway |
| Workforce infrastructure | No current sandboxing setup | Sandbox or code-execution tooling already in place |
| Auditability wants | Each discrete motion should be individually logged | Mixture final result issues greater than every inside step |
The Hybrid Actuality: Most Manufacturing Brokers Use Each
It’s value closing this out by pushing again gently on the framing of the article’s personal title. In observe, this isn’t a everlasting, once-and-for-all architectural choice — it’s a per-task judgment name, and Anthropic’s personal steerage treats it precisely that manner. Their advanced tool use release shipped Programmatic Device Calling alongside two companion options particularly meant to be layered collectively as wanted: a Device Search Device for locating the fitting software out of a giant library with out loading each definition upfront, and Device Use Examples for instructing a mannequin the conventions a schema alone can’t categorical. Their very own suggestion is to start out with whichever bottleneck is definitely limiting a given agent — context bloat from too many software definitions, massive intermediate outcomes polluting context, or parameter errors — and add the matching function, moderately than reaching for each functionality on day one.
A single well-built agent, in observe, tends to make use of plain software calling for its easy, single-shot lookups and swap to code execution the second a activity requires fan-out, aggregation, or dealing with information too massive or delicate to place in entrance of the mannequin instantly. The precise talent value constructing isn’t choosing a primitive as soon as — it’s recognizing, activity by activity, which one the work in entrance of you truly wants.
Conclusion
An motion primitive is infrastructure, not a choice, and the 2 examples constructed on this article show it with the identical fifteen strains of software definition beneath each. Get it proper and an agent handles a fan-out activity throughout fifteen cities — or two thousand expense line objects — in a single clear move. Get it unsuitable — attain for software calling on a activity that wants code execution — and nothing crashes. The agent nonetheless solutions. It simply does it slower, extra expensively, and with a context window quietly full of information no person truly wanted to learn.

