Abstract
This study quantifies gender and skin-tone bias in two widely deployed commercial image generators - Gemini Flash 2.5 Image (NanoBanana) and GPT Image 1.5 - to test the assumption that neutral prompts yield demographically neutral outputs. We generated 3,200 photorealistic images using four semantically neutral prompts. The analysis employed a rigorous pipeline combining hybrid color normalization, facial landmark masking, and perceptually uniform skin tone quantification using the Monk (MST), PERLA, and Fitzpatrick scales. Neutral prompts produced highly polarized defaults. Both models exhibited a strong "default white" bias (>96% of outputs). However, they diverged sharply on gender: Gemini favored female-presenting subjects, while GPT favored male-presenting subjects with lighter skin tones. This research provides a large-scale, comparative audit of state-of-the-art models using an illumination-aware colorimetric methodology, distinguishing aesthetic rendering from underlying pigmentation in synthetic imagery. The study demonstrates that neutral prompts function as diagnostic probes rather than neutral instructions. It offers a robust framework for auditing algorithmic visual culture and challenges the sociolinguistic assumption that unmarked language results in inclusive representation.
Abstract (translated)
这项研究量化了两种广泛部署的商业图像生成器——Gemini Flash 2.5 Image(NanoBanana)和GPT Image 1.5 中的性别和肤色偏见,以测试中立提示会产生人口统计学上中立输出这一假设。我们使用四个语义中性的提示生成了3,200张逼真的图像。分析采用了一种严谨的方法,结合混合色彩归一化、面部特征遮盖以及Monk(MST)、PERLA和Fitzpatrick肤色量表进行感知一致的肤色量化。 中立的提示产生了高度极化的默认设置。两个模型都表现出强烈的“默认白色”偏见(超过96% 的输出)。然而,在性别上,两者差异明显:Gemini 更倾向于生成女性形象;而GPT 则更倾向于生成男性形象,且皮肤色调较浅。这项研究提供了一个大规模、比较性的审计,使用了光照感知的色彩计量方法,将审美呈现与合成图像中的实际色素沉着区分开来。 该研究表明,中立提示作为诊断探针而非中立指令发挥作用。它为算法视觉文化的审核提供了稳健框架,并挑战了社会语言学假设,即无标记的语言会导致包容性的表现形式。
URL
https://arxiv.org/abs/2602.12133