Abstract
Large language models (LLMs) achieve impressive benchmark performance yet exhibit systematic gaps in realworld deployment contexts. This is particularly concerning for AI governance applications where LLMs might identify risk categories but fail to trace the causal mechanisms through which harms actually unfold. We propose that structuring documented AI risks as explicit causal patterns-sequential mechanisms, feedback loops, and quantified effects from empirical research can scaffold more reliable LLM reasoning through human-AI collaboration. We test this through a controlled study comparing pattern-augmented LLMs against vanilla LLMs using GPT-4o and Grok 4 to analyse five real estate AI deployment cases. Results show substantial improvements: GPT-4o improved in causal depth with Cohen’s d=1.20, while Grok showed even larger effects (d=1.93) alongside improved evidence grounding (d=2.10). Validation using Gemini 2.5 Pro and Deep Seek as judges confirmed findings with moderate-to-high inter-rater agreement (r=0.45 to 0.68). Qualitative analysis reveals explicit mechanistic transfer, with augmented outputs mapping patterns from domains such as aviation automation and reinforcement learning pricing to novel real estate contexts. These findings suggest that structured causal knowledge can enable a new modality of human-AI collaboration for AI risk assessment.
| Original language | English |
|---|---|
| Title of host publication | Augmenting Large Language Models with Causal Risk Patterns for AI Deployment Risk Assessment |
| Publisher | IEEE Xplore |
| Pages | 1-8 |
| ISBN (Electronic) | 979-8-3315-7330-0 |
| ISBN (Print) | 979-8-3315-7331-7 |
| DOIs | |
| Publication status | Published online - 22 Jul 2026 |
| Event | IEEE International Conference on AI and Data Analytics - Boston, United States Duration: 11 Jun 2026 → 12 Jun 2026 |
Conference
| Conference | IEEE International Conference on AI and Data Analytics |
|---|---|
| Abbreviated title | ICAD |
| Country/Territory | United States |
| City | Boston |
| Period | 11/06/26 → 12/06/26 |
Bibliographical note
© 2026 IEEE.Data Access Statement
The 100 experimental outputs, 87-pattern library, case studies, blinding key, raw judge scores from Gemini 2.5 Pro and DeepSeek, coding manual, and analysis code are available as open data at https://github.com/ggmcconomy/Causal RiskPatterns for AI Deployment Risk Assessment.
Funding
Kainos
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 3 Good Health and Well-being
-
SDG 9 Industry, Innovation, and Infrastructure
Keywords
- LLMs
- Risk Assessment
- AI
- AI risk assessment
- Large language models
- AI governance
- retrieval-augmented generation
- human-AI collaboration
- causal reasoning
Fingerprint
Dive into the research topics of 'Augmenting Large Language Models with Causal Risk Patterns for AI Deployment Risk Assessment'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver