ORCID

Abstract

Multi-agent AI systems, with minimal human oversight, are increasingly deployed in high-stakes domains like hiring, healthcare, and criminal justice. However, existing bias evaluation methodologies focus on isolated Large Language Model responses, failing to address sequential agent interactions where biases can propagate through decision chains undetected. Current approaches suffer from three limitations: absence of multi-stage bias tracking, inability to distinguish systematic discrimination from model non-determinism, and risk of contamination when agents infer bias evaluation intent. This study presents a novel evaluation framework incorporating demographic swapping methodology, contamination prevention architecture, and statistical analysis for multi-agent workflows. The framework employs dual presentation systems separating agent-visible content from research metadata and control scenarios to isolate bias from random variation. Evaluation across multiple models revealed unexpected preferences favouring traditionally disadvantaged groups, while 85% of apparent variation was attributable to model non-determinism rather than demographic factors. This approach advances methodological frameworks for tracking bias propagation in autonomous AI systems.

Keywords

Agentic Artificial Intelligence, Bias Detection, Ethical AI, Human-Centred AI, Non-determinism

Publication Date

2026-02-16

Event

3rd International Conference on Human-Centred AI - Education and Practice, HCAI-ep 2026

Publication Title

HCAI-ep 2026 - Proceedings of the 2026 Conference on Human Centered Artificial Intelligence - Education and Practice

Publisher

Association for Computing Machinery (ACM)

ISBN

9798400721533

First Page

46

Last Page

52

Deposit Date

2026-05-11

Creative Commons License

Creative Commons Attribution 4.0 International License
This work is licensed under a Creative Commons Attribution 4.0 International License.


Share

COinS