Creating safer reward functions for reinforcement learning agents in the gridworld

De Biase, Andres
Namgaudis, Mantas
Göteborgs universitet/Institutionen för data- och informationsteknikswe
University of Gothenburg/Department of Computer Science and Engineeringeng
2019-11-12T11:20:08Z
2019-11-12T11:20:08Z
2019-11-12
We adapted Goal-Oriented Action planning, a decision-making architecture common in video games into the machine learning world with the objective of creating a safer artificial intelligence. We evaluate it in randomly generated 2D grid-world scenarios and show that this adaptation can create a safer AI that also learns faster than conventional methods.sv
http://hdl.handle.net/2077/62445
engsv
Technology
Creating safer reward functions for reinforcement learning agents in the gridworldsv
text
Student essay
M2

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
gupea_2077_62445_1.pdf
Size:
612.22 KB
Format:
Adobe Portable Document Format
Description:
CSE Group 18 - De Biase & Namgaudis

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
876 B
Format:
Item-specific license agreed upon to submission
Description: