Author Identifier
Date of Award
2026
Keywords
explainable AI, mechanistic interpretability, concept-based explanations, causal interpretability, robotic vision, agentic interpretability
Document Type
Thesis
Publisher
Edith Cowan University
Degree Name
Doctor of Philosophy
School
School of Science
First Supervisor
Syed Afaq Ali Shah
0000-0003-2181-8445
Second Supervisor
Syed Mohammed Shamsul Islam
0000-0002-3200-2903
Third Supervisor
David Suter
0000-0001-6306-3023
Abstract
Intelligent agents have become integral to our daily tasks and interactive settings, operating across domains as diverse as robotic vision, multimodal reasoning, and mental health support. Despite their impressive performance gains, they remain largely black-box whose internal mechanisms are poorly understood. This opacity limits the ability to trust and deploy such systems responsibly in high-stakes settings. Understanding model behaviour means examining representational patterns, internal encodings, and their role in driving agent decisions. This thesis advances the interpretability of intelligent agents across three levels of increasing complexity, namely perceptual, representational, and behavioural. At the perceptual level, the thesis introduces a multimodal explainability framework for visual affordance learning in robotic vision, integrating visual attribution maps with language-based explanations to communicate not only where an agent attends but also what that attention implies for action. To address the challenge of evaluating such explanations at scale, the thesis further introduces DXAI, a large-scale synthetic dataset of 100,000 image-heatmap pairs across 100 object categories, enabling systematic bench marking of visual attribution methods against ground-truth explanation maps. At the representational level, the thesis proposes an automated framework for concept-based interpretability in vision-language models. Rather than relying on predefined concept vocabularies, the framework constructs data-grounded concept sets, filters them for visual presence, and aligns them with sparse internal features discovered by sparse autoencoders. At the behavioural level, the thesis examines how human-aligned behaviour is encoded and causally sustained in therapeutic conversational agents. Our proposed agentic M-AIDE framework reveals that empathic behaviour is encoded in structured internal features aligned with various dimensions of empathy, which strengthen with model depth. Building on this, we propose another framework, called Lezak, that applies layer-wise causal interventions to test whether sparse internal features are causally necessary for sustaining therapeutic behaviour, uncovering that different layers contribute distinct functional roles. These contributions provide a foundation for developing interpretable intelligent agents that support their responsible deployment in high-stakes settings for images and text data.
Access Note
Access to this thesis is embargoed until 23rd September 2027
Recommended Citation
Mirnateghi, N. (2026). Towards transparent artificial intelligence: Unravelling the black-box behaviour of intelligent agents. Edith Cowan University. https://doi.org/10.25958/qqx4-bm12