An Outcome Reward Model (ORM) is a feedback mechanism that scores only the final response generated by a model, without evaluating the correctness of intermediate reasoning steps. It is simpler to train but less granular than step-by-step reward models.
Helps AI builders design and scale robust architectures; mastering the implementation of Outcome Reward Model improves latency, accuracy, and operational efficiency for basic classification, text summarization, and simple question-answering validation.
Outcome Reward Models (ORMs) are feedback systems that score only the final correctness or quality of a model's complete response. While ORMs are easy to configure and require less annotation effort than process-based models, they provide less guidance during multi-step reasoning. Without step-level feedback, ORMs can inadvertently reward model outputs that reach correct conclusions through incorrect or hallucinated logical steps.
ORMs are much easier and cheaper to train because labeling only the final correctness of a response is faster than labeling every reasoning step.
It can reward "logical alignment by coincidence" where a model arrives at the correct answer through flawed logic or guessing.
Reference this definition in your articles, research, or documentation to credit this source:
We currently have no direct coverage articles matching "Outcome Reward Model". Explore trending global AI topics below instead.
Across Scottish Water's Capital Investment (CI) programme, teams need fast answers...
Databricks Inc. today announced that it has raised $5 billion in funding at a $190 billion valuation. The round was led by Coatue, Blackstone, MGX, T. Rowe Price and Sixth Street Growth. The funds were joined by more than a half-dozen other backers, most of which are returning investors. The...
AI is expensive, Ali Ghodsi tells TechCrunch. With so many investors wanting into his latest round, he said yes to more than planned.
Google is rolling out Gemini 3.7 Flash , a new version of its workhorse AI model that puts coding, agentic workflows and knowledge work at the center of the...