Rejection Sampling (in LLMs) is a data curation technique where a generator model produces multiple candidate answers, and a separate evaluator model filters out low-quality outputs. The remaining high-quality responses are then used for supervised fine-tuning.
Helps AI builders design and scale robust architectures; mastering the implementation of Rejection Sampling improves latency, accuracy, and operational efficiency for high-quality instruction dataset creation, code correctness filtering, and model bootstrapping.
Rejection sampling is a dataset purification technique widely used in alignment training (SFT and RLHF). A source model generates multiple candidate completions for a set of instructions, and a separate evaluator (or reward model) filters out low-scoring or incorrect responses. The remaining top-quality examples are saved to form clean demonstration datasets, bootstrapping model capabilities without human labeling.
It filters out poor reasoning steps or incorrect outputs, ensuring the model only learns from high-quality, correct demonstrations.
It is often referred to as best-of-N sampling or self-training with selection.
Reference this definition in your articles, research, or documentation to credit this source:
We currently have no direct coverage articles matching "Rejection Sampling". Explore trending global AI topics below instead.
Prompt engineering in Amazon Quick shapes how accurately its AI-powered feature respond to your requests. Part 1 of a two-part series covers the...
Part 2 of our Amazon Quick prompt engineering series goes component by component. Learn the prompt patterns that get the best results from Amazon Quick...