Date of Award

2026

Keywords

explainable AI, mechanistic interpretability, concept-based explanations, causal interpretability, robotic vision, agentic interpretability

Document Type

Thesis

Publisher

Edith Cowan University

Degree Name

Doctor of Philosophy

School

School of Science

First Supervisor

Syed Afaq Ali Shah ORCID iD 0000-0003-2181-8445

Second Supervisor

Syed Mohammed Shamsul Islam ORCID iD 0000-0002-3200-2903

Third Supervisor

David Suter ORCID iD 0000-0001-6306-3023

Abstract

Intelligent agents have become integral to our daily tasks and interactive settings, operating across domains as diverse as robotic vision, multimodal reasoning, and mental health support. Despite their impressive performance gains, they remain largely black-box whose internal mechanisms are poorly understood. This opacity limits the ability to trust and deploy such systems responsibly in high-stakes settings. Understanding model behaviour means examining representational patterns, internal encodings, and their role in driving agent decisions. This thesis advances the interpretability of intelligent agents across three levels of increasing complexity, namely perceptual, representational, and behavioural. At the perceptual level, the thesis introduces a multimodal explainability framework for visual affordance learning in robotic vision, integrating visual attribution maps with language-based explanations to communicate not only where an agent attends but also what that attention implies for action. To address the challenge of evaluating such explanations at scale, the thesis further introduces DXAI, a large-scale synthetic dataset of 100,000 image-heatmap pairs across 100 object categories, enabling systematic bench marking of visual attribution methods against ground-truth explanation maps. At the representational level, the thesis proposes an automated framework for concept-based interpretability in vision-language models. Rather than relying on predefined concept vocabularies, the framework constructs data-grounded concept sets, filters them for visual presence, and aligns them with sparse internal features discovered by sparse autoencoders. At the behavioural level, the thesis examines how human-aligned behaviour is encoded and causally sustained in therapeutic conversational agents. Our proposed agentic M-AIDE framework reveals that empathic behaviour is encoded in structured internal features aligned with various dimensions of empathy, which strengthen with model depth. Building on this, we propose another framework, called Lezak, that applies layer-wise causal interventions to test whether sparse internal features are causally necessary for sustaining therapeutic behaviour, uncovering that different layers contribute distinct functional roles. These contributions provide a foundation for developing interpretable intelligent agents that support their responsible deployment in high-stakes settings for images and text data.

Access Note

Access to this thesis is embargoed until 23rd September 2027 

Available for download on Thursday, September 23, 2027

Share

 
COinS
 

Link to publisher version (DOI)

10.25958/qqx4-bm12