An Agentic Framework for Real-Time Root Cause Analysis in Kubernetes via the Model Context Protocol - Designing and Evaluating an MCP-Grounded LLM Agent for Kubernetes Incident Triage in a ChatOps Environment

Loading...
Thumbnail Image

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Root cause analysis in Kubernetes incident triage is fragmented across observability tools and operator expertise, and large language models proposed as diagnostic assistants typically lack a standardised mechanism for accessing live system state. This thesis investigates whether the Model Context Protocol (MCP) can be used to ground a commercial language model in live Kubernetes telemetry, in collaboration with an industry partner referred to as Company X. Following a Design Science Research process across three iteration cycles, the resulting artifact we developed, ARGUS, combines Claude Sonnet 4.6 with MCP-mediated read-only access to Kubernetes, Prometheus, Loki, and NATS, retrieval-augmented few-shot context from past incidents, and an eARCO-structured prompt that drives a ReAct investigation loop. The output is a structured RCA summary posted as a Slack thread reply to the incoming alert. ARGUS was evaluated on a dedicated test cluster through three complementary methods: tool-trace analysis across ten controlled fault injection scenarios, a structured rubric scored independently by two raters and a third-party adjudicator, and semi-structured interviews with six on-call engineers. The agent retrieved sufficient evidence to name the correct root cause in every scenario, with all observed tool-call failures originating in the MCP infrastructure layer rather than in the agent’s reasoning. Rubric scores were strong on Correctness and Evidence backing and showed a clearly bounded weakness on Actionability, the only dimension on which any Not met cell was returned. The interview themes converged on the same asymmetry: the diagnostic core of the output was trusted but the Recommended Fixes block was consistently questioned, and the tool’s value was largest where the responder was least familiar with the affected subsystem or incident alert. The thesis contributes a working composition of existing components into an MCP grounded RCA assistant deployed at an industrial site, a three-method evaluation design that separates evidence retrieval from reasoning quality and practitioner perception, and the load-bearing finding that the diagnostic reliability of such agents sits well ahead of their prescriptive reliability.

Description

Keywords

Root cause analysis, Kubernetes, large language models, MCP, incident response

Citation

ISBN

Articles

Department

Defence location

Collections

Endorsement

Review

Supplemented By

Referenced By