Home › Glossary › What Is Reinforcement Learning from Human Feedback (RLHF)?

What Is Reinforcement Learning from Human Feedback (RLHF)?

Reinforcement Learning from Human Feedback (RLHF) is a machine learning technique that fine-tunes a model's behavior using a reward signal derived from human preference judgments, rather than a fixed labeled dataset alone.

How RLHF works

Human reviewers rank or rate multiple model outputs for the same input. Those preference judgments train a reward model, which is then used to further optimize the underlying model’s behavior through reinforcement learning.

Why it’s relevant to data providers

RLHF pipelines depend on structured human feedback data collection and management, which shares infrastructure needs with other AI training data workflows: annotation, quality control, and scalable data delivery.