Metanormative Theory Informs RL-Based Moral Agent Design
A new paper on arXiv explores how metanormative theory can inform the design of moral and value-aligned artificial agents using reinforcement learning (RL). The paper addresses the growing trend in machine ethics and value alignment to use RL, which has sidelined philosophical literature. The authors aim to draw on recent metanormative theory to provide clearer criteria for classifying RL agent behavior as moral and to evaluate and compare RL-based approaches. The paper is categorized under Computer Science > Artificial Intelligence and was submitted to arXiv with the identifier 2608.08220. It discusses the overlap between machine ethics and value alignment, focusing on designing agents that align with human values and act ethically. The paper's two main goals are to extract useful ideas from metanormative theory and to examine RL architecture through that lens. The submission includes references, citations, and tools, and is part of the arXivLabs framework, which supports community collaborations. The paper emphasizes the importance of philosophical insights in the development of ethical AI, countering the current trend that prioritizes technical methods over philosophical foundations.
Key facts
- Paper titled 'Metanormative Theory for RL-Based Moral Agents' on arXiv.
- Filed under Computer Science > Artificial Intelligence.
- Addresses machine ethics and value alignment.
- Focuses on using reinforcement learning (RL) for designing moral agents.
- Draws on metanormative theory to inform RL agent design.
- Aims to provide criteria for classifying RL behavior as moral.
- Paper ID: 2608.08220.
- Part of arXivLabs framework for community collaborations.
Entities
Institutions
- arXiv