---
type: "GlossaryTerm"
title: "RLHF (Reinforcement Learning from Human Feedback)"
description: "The dominant method for aligning LLMs with human preferences. Humans rate model outputs, and the model is trained to prefer higher-rated answers. Can lead to Mo"
resource: "https://www.contextstudios.ai/glossary/rlhf"
language: "en"
tags: ["engineering"]
generated:
  by: "process:contextstudios-md/1"
  at: "2026-10-08T20:45:30.958Z"
status: "stable"
---

# RLHF (Reinforcement Learning from Human Feedback)

The dominant method for aligning LLMs with human preferences. Humans rate model outputs, and the model is trained to prefer higher-rated answers. Can lead to Mode Collapse as 'typical' answers are systematically preferred.

## Business Value



RLHF is how models like ChatGPT and Claude become helpful and safe. Understanding its mechanics helps you predict model behavior and work around its limitations.

## Context Studios Perspective



RLHF is powerful but imperfect. We help clients understand where RLHF-induced behaviors help or hinder their use cases – and how to prompt around limitations.
