Learning Morality from Games? An Analysis of AI-Learned Action Preferences in Annotated Scenarios from Detroit: Become Human
Date
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
This thesis investigates whether an AI agent can learn morally relevant decision patterns from narrative-driven game scenarios. The project uses selected scenarios from Detroit: Become Human, a game built around authored choices, branching consequences, and morally charged situations. These scenarios were reconstructed as structured text-based interactive environments in which an agent observes a written narrative situation, chooses among available actions, and receives feedback based on moral annotations and final outcomes. The scenarios were annotated using a framework that considers duties, consequences, and effects on both the acting character and others. Outcome values were also in formed by human survey responses. An adapted SAC-based reinforcement-learning framework was then used to train agents with different moral-reward emphases and compare them with an untrained version of the same architecture. Evaluation combined deterministic scenario rollouts with controlled action-label probes to inspect selected actions, policy scores, critic values, and ranking changes. The results show that training produced stable and interpretable action preferences. Compared with the untrained agent, trained agents developed clearer preferences, more consistent scenario behaviour, and stronger separation between actions they tended to favour or avoid. Differences between reward settings were especially visible in morally mixed cases, such as actions involving coercion, reassurance, self-sacrifice, or outcome-focused success. However, the findings do not show that the agents learned morality in a human-like sense. Their behaviour remained shaped by the reward design, the reconstructed scenario structure, and the wording of action labels. The agents also showed limited transfer to differently phrased actions. The thesis, therefore, concludes that AI agents can learn morally relevant, reward-shaped action preferences from annotated narrative-game scenarios, but these preferences should be understood as patterns learned within a controlled textual representation, not as general moral understanding.