Glossary · term

Test-time RL

A method (TTRL) that applies reinforcement learning to unlabeled data at inference time, allowing a model to improve itself without ground-truth labels. It uses majority voting across multiple responses as a reward signal. Work by Yuxin Zuo et al. (Tsinghua/Shanghai AI Lab, April 2025).

Training2025–2026Wave 2 · 2024Maturity: 1/5

Maturity rationale

single R2 source, 2025-26 neologism

References

Author: Społeczność / Anonimowi