Creating safer reward functions for reinforcement learning agents in the gridworld
| De Biase, Andres | ||
| Namgaudis, Mantas | ||
| Göteborgs universitet/Institutionen för data- och informationsteknik | swe | |
| University of Gothenburg/Department of Computer Science and Engineering | eng | |
| 2019-11-12T11:20:08Z | ||
| 2019-11-12T11:20:08Z | ||
| 2019-11-12 | ||
| We adapted Goal-Oriented Action planning, a decision-making architecture common in video games into the machine learning world with the objective of creating a safer artificial intelligence. We evaluate it in randomly generated 2D grid-world scenarios and show that this adaptation can create a safer AI that also learns faster than conventional methods. | sv | |
| http://hdl.handle.net/2077/62445 | ||
| eng | sv | |
| Technology | ||
| Creating safer reward functions for reinforcement learning agents in the gridworld | sv | |
| text | ||
| Student essay | ||
| M2 |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- gupea_2077_62445_1.pdf
- Size:
- 612.22 KB
- Format:
- Adobe Portable Document Format
- Description:
- CSE Group 18 - De Biase & Namgaudis
License bundle
1 - 1 of 1
Loading...
- Name:
- license.txt
- Size:
- 876 B
- Format:
- Item-specific license agreed upon to submission
- Description: