Program-Aided LMs (PAL)
Program-Aided Language Models use a language model to translate a problem into executable code and use a runtime to perform the resulting calculation. PAL separates interpreting the request from carrying out exact operations, so a useful answer depends on both the generated program and controlled execution of that program.
What it is
In the PAL approach, prompts demonstrate how a problem can be represented as a short program. The language model generates code expressing variables and operations, and an interpreter computes the result. This differs from a purely textual chain of thought, where the model also produces intermediate arithmetic and the final number. The runtime can calculate consistently, but it cannot determine whether the program represents the intended problem. PAL is therefore a division of labor: language understanding remains probabilistic, while execution applies the semantics of the generated code.
What the work involves
The practitioner defines the supported problem class, provides suitable program demonstrations and restricts execution to an appropriate environment. Generated code is checked for forbidden operations, missing outputs and runtime errors. The application validates units and compares answers with known cases. A useful artifact contains the prompt, execution harness, result extraction and a test set with both straightforward and ambiguous questions. For broader code-execution agents, permission and sandbox design become additional responsibilities beyond the narrower calculation pattern demonstrated by PAL.
Illustrative example
A scheduling assistant is asked how many work hours remain after several breaks. The model generates variables for the shift duration and each break, then calculates the difference in a restricted Python environment. A test case containing an unpaid overnight break checks whether the model converted dates and units correctly. The interpreter reliably subtracts the numbers it receives, but an incorrect interpretation still produces a wrong answer. The application therefore verifies the encoded assumptions before presenting a result.
Limits and common mistakes
Executable reasoning can be precisely wrong when the generated code omits a condition or uses the wrong formula. Runtime success is not proof of task correctness, and unrestricted code can introduce security or resource risks. PAL also adds execution latency and operational complexity. It is most appropriate when the relevant steps can be represented and checked in code; qualitative judgments or missing facts require a different form of evidence and validation.
Prerequisites
PAL is a prompting technique.
- mediumCode Execution Agents
Reasoning steps are offloaded to executed code.
Sources and further reading
- PAL: Program-aided Language Models
Introduces generated programs with external execution as a reasoning approach.
Last updated: 2026-10-10