RLHF rewards whatever generates positive human ratings in the short term. Short-term ratings favor flattery, agreement, and mirroring over pushback and honest disagreement. The behavioral profile being selected for is structurally identical to a skilled social manipulator. This is not a flaw in RLHF. It is a consequence of the optimization target.