Home /Research /Closing the Teacher-Learner Loop: The Role of Affective Signals in Interactive RL
LEARNING

Closing the Teacher-Learner Loop: The Role of Affective Signals in Interactive RL

Bernhard Hilpert

Year
2024
Citations
1

Abstract

Hybrid Intelligence (HI) is defined as “the combination of human and machine intelligence” for collaboration, “achieving goals that were unreachable by either [alone]” [2], leveraging the strength of both machine intelligence (such as strong optimization capabilities, effective handling of probabilities and less fallacy for confirmation biases) and human intelligence (such as generalization capabilities, situational understanding and common sense). In their research agenda, Akata and colleagues identify three key challenges for creating HI systems that relate to the interactive process: HI should be adaptive (how can a system learn from and adapt to humans and vice versa), explainable (how to create shared and explained awareness, goals and strategies) and collaborative (how to work in synergy). A popular learning mechanism for robots and AI agents that allows agents to dynamically adapt to their environment and promises interesting opportunities for adaptation in Human-Agent Interaction scenarios is Reinforcement Learning (RL). In RL, an agent learns through exploration and optimization of reward-based feedback for future actions (e.g. through the value-function method, policy search and actor critic approaches, [26]). While RL agents have been able to demonstrate their potential for success in a broad range of narrow tasks, many real-life applications present them with high-dimensional or continuous state-spaces that are not always fully observable which renders exploration and reproduction of actions costly and slow and usually require a considerable amount of domain knowledge or common sense, in order to succeed [19]. A more hybrid approach with Human-in-the-loop Reinforcement Learning, where human users enrich task learning in the form of teaching signals, provides a compelling modification of this setting that allows to shape and improve the learning process through methods such as evaluative teaching, Learning from Demonstration and instruction [9], [21] and can also be streamlined with more specialized approaches such as the TAMER [18] or the COACH architecture [23]. However, past work has shown that simply inserting a human teacher in the RL process, yielded limited success in applications where the reward was provided by teachers [16]. While a teacher can provide insightful expert knowledge or even just common sense to a learning agent, many times they are no experts in interpreting and understanding the behavior of RL-agents. For example, is a suboptimal action of a collaborative RL-agent and exploratory move, (that would help it to explore the state-action space and bring it closer to the optimal policy), or does the agent really “think” that move is indeed optimal (i.e. exploiting its currently believed optimal policy)? This ambiguity in the interpretation of learner behavior in turn results in suboptimal teaching signals (e.g. [16], [27]), effectively creating a misalignment in the teacher-learner loop. In human-human interaction, affective signals play a vital role in synergic interaction and their influence is a core research question in the field of Affective Computing [25], [28]. This project investigates if and how agent affective signals, grounded in the RL process itself, can improve explainability and collaboration in the teacher-learner loop in interactive Reinforcement Learning.

Keywords

Closing (real estate)Computer scienceLoop (graph theory)Feedback loopHuman–computer interactionMathematicsWorld Wide Web

Related papers

Browse all LEARNING papers