dorsal/arxiv
View SchemaArchitecture inside the mirage: evaluating generative image models on architectural style, elements, and typologies
| Authors | Jamie Magrill, Leah Gornstein, Sandra Seekins, Barry Magrill |
|---|---|
| Categories | |
| ArXiv ID | 2601.09169vv1 |
| URL | https://arxiv.org/abs/2601.09169 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Generative artificial intelligence (GenAI) text-to-image systems are increasingly used to generate architectural imagery, yet their capacity to reproduce accurate images in a historically rule-bound field remains poorly characterized. We evaluated five widely used GenAI image platforms (Adobe Firefly, DALL-E 3, Google Imagen 3, Microsoft Image Generator, and Midjourney) using 30 architectural prompts spanning styles, typologies, and codified elements. Each prompt-generator pair produced four images (n = 600 images total). Two architectural historians independently scored each image for accuracy against predefined criteria, resolving disagreements by consensus. Set-level performance was summarized as zero to four accurate images per four-image set. Image output from Common prompts was 2.7-fold more accurate than from Rare prompts (p < 0.05). Across platforms, overall accuracy was limited (highest accuracy score 52 percent; lowest 32 percent; mean 42 percent). All-correct (4 out of 4) outcomes were similar across platforms. By contrast, all-incorrect (0 out of 4) outcomes varied substantially, with Imagen 3 exhibiting the fewest failures and Microsoft Image Generator exhibiting the highest number of failures. Qualitative review of the image dataset identified recurring patterns including over-embellishment, confusion between medieval styles and their later revivals, and misrepresentation of descriptive prompts (for example, egg-and-dart, banded column, pendentive). These findings support the need for visible labeling of GenAI synthetic content, provenance standards for future training datasets, and cautious educational use of GenAI architectural imagery.
{
"annotation_id": "cc372059-bcbf-4f01-b009-d47ce3477212",
"date_created": "2026-02-17T05:53:20.072000Z",
"date_modified": "2026-02-17T05:53:20.072000Z",
"file_hash": "6c220467bf45812d51bfe90a0de6a5c3d790f23156dad0979fee63a2fa5cb730",
"private": false,
"record": {
"abstract": "Generative artificial intelligence (GenAI) text-to-image systems are increasingly used to generate architectural imagery, yet their capacity to reproduce accurate images in a historically rule-bound field remains poorly characterized. We evaluated five widely used GenAI image platforms (Adobe Firefly, DALL-E 3, Google Imagen 3, Microsoft Image Generator, and Midjourney) using 30 architectural prompts spanning styles, typologies, and codified elements. Each prompt-generator pair produced four images (n = 600 images total). Two architectural historians independently scored each image for accuracy against predefined criteria, resolving disagreements by consensus. Set-level performance was summarized as zero to four accurate images per four-image set. Image output from Common prompts was 2.7-fold more accurate than from Rare prompts (p \u003c 0.05). Across platforms, overall accuracy was limited (highest accuracy score 52 percent; lowest 32 percent; mean 42 percent). All-correct (4 out of 4) outcomes were similar across platforms. By contrast, all-incorrect (0 out of 4) outcomes varied substantially, with Imagen 3 exhibiting the fewest failures and Microsoft Image Generator exhibiting the highest number of failures. Qualitative review of the image dataset identified recurring patterns including over-embellishment, confusion between medieval styles and their later revivals, and misrepresentation of descriptive prompts (for example, egg-and-dart, banded column, pendentive). These findings support the need for visible labeling of GenAI synthetic content, provenance standards for future training datasets, and cautious educational use of GenAI architectural imagery.",
"arxiv_id": "2601.09169",
"authors": [
"Jamie Magrill",
"Leah Gornstein",
"Sandra Seekins",
"Barry Magrill"
],
"categories": [
"cs.CV",
"cs.CY"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Architecture inside the mirage: evaluating generative image models on architectural style, elements, and typologies",
"url": "https://arxiv.org/abs/2601.09169",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "2a39381c-a6a8-45a3-bf08-f32e3b01067e",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}