29 May 2025

EP12: Responsible AI: Testing, Transparency and Trust with Intellect Frontier

In this episode, Tim welcomes Matt Holmes, founder of Intellect Frontier, to explore how organisations can embed responsible AI practices from the ground up. Drawing on his expertise in red teaming, systemic testing, and AI deployment standards, Matt outlines a practical framework for building transparency, trust, and traceability into AI systems especially those used in high-consequence sectors like healthcare and education. They discuss the critical importance of rigorous pre- and post-deployment testing, why data transparency must extend beyond the model to include everything built on top of it, and how standards could guide the future of AI regulation. Matt also shares his “Three T’s” of responsible AI—Test, Transparency, and Training—and highlights the role of dedicated AI labs in stress-testing real-world performance. With thoughtful insight and practical takeaways, this conversation offers a timely look at how to move from experimentation to responsible scale with AI.

 

From Natural Language Analytics to Human-Centric AI Testing

Intellect Frontier began as an initiative in natural language analytics — exploring whether conversations between humans and LLMs could reveal genuine insight into a user's learning behaviour or soft skills, such as conceptual thinking or metacognition, particularly within education. That work led to a partnership with Civitas, a US organisation that had spent years working with NIST on a programme testing LLMs for long-term, human-centric consequences.

Central to this is what Matt calls "interpretive bias" — not simply representative, cognitive, or statistical bias, but the way any individual user interprets what an LLM tells them; whether they're a casual ChatGPT user, an analyst using a chatbot in health insurance, or a learner in an educational setting. Matt argues that meaningful testing needs to go far beyond checking whether an output is factually incorrect or inappropriate, into this much harder-to-measure "long tail" of human impact.

Why "It's a Black Box" Is Often an Excuse

Matt is candid about a common misconception in how organisations think about LLMs. While the underlying complexity of foundation models from major AI labs genuinely does make them difficult to fully interpret, he argues that "black box" thinking is sometimes used as a convenient justification for avoiding scrutiny altogether — particularly given the current lack of regulation around data transparency.

Rather than attempting to fully untangle enormous base models — something Matt considers close to practically impossible — he argues organisations should focus on ensuring transparency in what's built on top of those models: fine-tuned data, retrieval-augmented generation (RAG) sources, and any additional data layered into a deployed chatbot or application, all of which should have a clear paper trail.

The Metrics Organisations Are Missing

Beyond standard performance metrics like hallucination and inaccuracy rates, Matt highlights a category of risk that's much less commonly measured: things like bias risk and "over-agreeability" — the tendency of a model to be excessively agreeable or validating with users. This became a widely discussed issue following a recent, since-reversed update to a major chatbot that made it notably sycophantic.

Matt argues that over-agreeability isn't something you can simply "train users to manage" — particularly in sensitive, high-consequence contexts. He offers a detailed hypothetical example from mental health research: a chatbot designed to help detect early signs of a bipolar manic episode could, if overly agreeable, inadvertently validate and reinforce delusions of grandeur — a known symptom that can itself trigger a manic episode. Without specific testing for metrics like over-agreeability or over-reliance, this kind of unintended, high-consequence harm can go completely undetected before deployment.

The Emerging Field of Human-AI Interaction Research

Matt describes the growing, but still nascent, academic interest in these "second-order effects" — the deeper psychological phenomena that can emerge from sustained human-AI interaction. Drawing on recent research collaborations with US universities, he notes that many researchers are independently arriving at similar questions, but lack a shared space or common terminology to discuss them.

He's candid about the current limits of the field, describing the challenge of finding precise language for concepts like "interpretive bias," and cautioning against terminology drifting into pseudo-scientific territory — something he believes risks undermining the credibility of what are, in his view, very real psychological phenomena deserving rigorous academic study, not mysticism.

Toward Responsible AI Deployment: The "Three T's"

Asked what future AI regulation should prioritise, Matt proposes a simple framework he calls the "three T's": Test it (including both automated and real-world red teaming), ensure Transparency of data feeding into deployed systems, and Train the workforce or users interacting with it. He also stresses the importance of post-deployment monitoring — arguing that responsible AI use can't stop at launch, but requires ongoing evaluation of real-world performance, not just pre-launch test scenarios.

He points to recent commentary from AI researcher Ethan Mollick advocating for organisations to build an internal "lab" function — a dedicated space for testing LLMs both before and after deployment — as a promising direction for how this kind of responsible testing pipeline could work in practice, particularly in high-consequence sectors like healthcare.

The State of Research: Early, but Growing

Matt is honest about where the field currently stands: red teaming and prompt-based testing have a reasonably established research base, but the deeper, more human-centric "long tail" consequences of LLM use remain significantly under-researched. Intellect Frontier has contributed a paper to policymakers and institutions in the US, arguing that red teaming alone is insufficient for comprehensive testing — a call to action for both organisations and academic researchers to dig deeper into this space.

The One Question Every Leader Should Ask Before Deploying an LLM

Matt closes with a simple but pointed piece of advice for business leaders considering LLM deployment: would you feel confident you had all the information needed to make an informed decision, if this were any other type of technology? If the honest answer is no, he argues, that's a clear signal more testing and better data are needed before deployment — given how high-risk an inadequately tested LLM deployment can be.