This is perhaps our most practically significant conclusion, since generative AI appears to hold great promise for many applications requiring fine-grained moral judgment and decision-making at society scale, and RLHF has emerged as the standard approach for aligning generative systems with human values.
← all excerpts
Moral disagreement and the limits of AI value alignment: <i>a dual challenge of epistemic justification and political legitimacy</i>.
1
—
—