Glossary · term

Tool-Integrated Reinforcement Learning / TIR-RL

RL for reasoning models in which thinking steps are coupled with tool calls (code, computation, search), moving Tool-Integrated Reasoning from prompting into post-training. The SimpleTIR paper stabilizes multi-step training by removing trajectories with void turns (steps containing neither code nor an answer) from the policy update while keeping them in the advantage estimation; starting from a Qwen2.5-7B base it reaches 50.5 on AIME24. Xue, Zheng, Liu, et al., ICLR 2026.

Training2026Wave 3 · 2025–26Maturity: 1/5

Maturity rationale

speculative / early neologism (warning)

References

Author: Społeczność / Anonimowi