CognArtive — LLMs for Automating Art Analysis
CognArtive
Anonymous · under review 15,000+ artworks read by machine

CognArtive

Large Language Models for Automating Art Analysis and Decoding Aesthetic Elements

A structured art analysis protocol applied at collection scale, and a case study of what happens when part of aesthetic interpretation is delegated to AI.

Explore 20 interactive figures See the protocol
chiaroscuro · impasto · sfumato · glazing · crosshatching · stippling · pointillism · diffused light · dappled light · monochromatic · earthy · pastel ·  chiaroscuro · impasto · sfumato · glazing · crosshatching · stippling · pointillism · diffused light · dappled light · monochromatic · earthy · pastel · 
nostalgia · serenity · melancholy · dramatic expression · hope · grief · passion · joy · despair · anxiety · calm · energy ·  nostalgia · serenity · melancholy · dramatic expression · hope · grief · passion · joy · despair · anxiety · calm · energy · 
15,000+
Artworks
23
Artists
34
Styles
8
Questions
4
Embedding models
4.33/5
Human rating
01

Hi, and welcome

You are in the right place. This page is the companion to the paper: the abstract, the method, what we found, and every interactive figure from the supplementary material, ready to explore.

Abstract

Multimodal language models can increasingly describe both the visual and expressive qualities of artworks, raising questions about how aesthetic interpretation changes when part of the process is delegated to AI. We present CognArtive, a framework that applies a structured art analysis protocol to more than 15,000 artworks from 23 artists and 34 styles. GPT vision generates artwork level interpretations across dimensions including form, color, light, movement, material, technique, temporal cues, figures, and emotional expression, while GPT and Gemini support synthesis and collection scale analysis. We evaluate how these model generated descriptions align semantically with established art style descriptions using four embedding models and find substantial variation across aesthetic dimensions. Beyond large scale analysis, CognArtive provides a case study of interpretive agency in creative AI, showing how human selected criteria, model representations, canonical categories, and corpus composition jointly shape computational interpretations of art.

02

The pipeline

Human selected questions define the analytical dimensions, GPT 4V interprets the visual evidence, and GPT 4 with Gemini 2.0 synthesize the outputs for collection scale analysis, semantic evaluation, and interactive visualization.

STAGE 01

Artwork corpus

More than 15,000 digitized artworks from WikiArt: 23 artists, 34 recorded styles, roughly the fifteenth to the twenty first century.

STAGE 02

Eight questions

A human defined framework derived from formal art analysis. Seven questions follow Hodge (2024); the eighth identifies the presence and type of figures.

STAGE 03

GPT 4V interpretation

Each artwork image and the eight questions go to GPT 4V, which produces descriptions of the corresponding aesthetic dimensions.

STAGE 04

Synthesis

Responses are cleaned, then processed with GPT 4 and Gemini 2.0 into structured representations for qualitative reading and quantitative aggregation.

STAGE 05

Three outputs

Collection scale analysis across artists, styles and periods; semantic comparison with art style descriptions; and the interactive figures on this page.

03

Eight questions, one artwork at a time

Every artwork is put to the same eight questions. Hover a spoke to hold it, and the panel shows what the model returns for that dimension, what it found across the collection, and how well it agrees with art historical language.

{{ qRoman }}
{{ qName }}
{{ qCount }}
Question {{ qRoman }} of VIII{{ qName }}

{{ qDesc }}

{{ qFinding }}

Style alignment{{ qAlignLabel }}
{{ qAlignNote }}
Human rating{{ qRatingLabel }}
Mean over 100 artworks, three annotators
Vocabulary returned
{{ t }}
More than twenty recurring emotional themes

The vocabulary the model reaches for when it reads feeling

Positive Neutral Negative Positive and neutral interpretations occur more frequently than negative ones. The grouping shown here is illustrative; measured counts are in Figures S6 and S7.
Tonal register across six historical periods

Twelve tones

{{ t.roman }}
{{ t.name }}
{{ t.hex }}
{{ t.name }}

{{ t.desc }}

{{ t.trend }}

Twelve tonal categories recur across the six periods. Hover a band for its register. Swatches are illustrative; the measured distributions are in Figure S5.

04

What we found

The patterns below describe model generated annotations within the sampled corpus. Each card opens the interactive figure behind it.

Form and scale

Natural forms remain prominent, while geometric and irregular forms become more frequent in later samples; reduced scales also become more visible over time.

Color

Color distributions show broad temporal variation, with monochromatic tones stronger in earlier samples and muted tones more visible later.

Emotion

CognArtive identifies more than twenty emotional themes, with positive and neutral interpretations occurring more frequently than negative ones.

Light and contrast

Light and contrast are frequently associated with highlights, depth, and texture, while contrast, shadow, and chiaroscuro are common visual effects. Diffused and soft lighting appear frequently, and emphasis or directing attention are common inferred purposes across major styles.

Movement

Conveyed and implied movement occur more frequently than literal movement.

Material and technique

Oil and canvas dominate material descriptions, while other media show distinct temporal patterns. Blending and layering are among the most frequent techniques, alongside variations in brushwork, line, crosshatching, scraping, and pointillism.

Time and figures

Morning and afternoon are common temporal cues, while human figures dominate the represented subjects.

Human evaluation

Three annotators rated 100 randomly sampled artworks across ten criteria: a mean of 4.33 out of 5, with 87.4% of 3,000 ratings scored 4 or 5.

See the full ratings →
05

Semantic alignment

Median cosine similarity between CognArtive descriptions and established art style descriptions, across four embedding models. Techniques show the strongest aggregate alignment, followed by Movement and Material.

{{ readTag }}
{{ readTitle }}

{{ readNote }}

0.37 0.70
{{ m.short }}
Aggregate
{{ row.name }}
{{ c.value }}
{{ row.agg }}

Aggregate reports the median after aggregation across embedding models. BGE M3 and NV Embed v2 return consistently higher similarities than Sentence BERT and OpenAI for every dimension, so the choice of embedding space shifts the absolute values more than the ranking between dimensions.

Human evaluation

Ten questions, three annotators

Each annotator rated all 100 randomly sampled artworks on a five point Likert scale, producing 3,000 individual ratings.

4.33/5
Overall mean
87.4%
Rated 4 or 5
Q1

The description of form and scale accurately matches what I see in the artwork.

4.37SD 0.69
90.0% rated 4 or 5
Q2

The description of colors accurately matches the colors I see in the artwork.

4.33SD 0.72
88.7% rated 4 or 5
Q3

The description of light and contrast accurately reflects the lighting and contrast in the artwork.

4.35SD 0.71
88.7% rated 4 or 5
Q4 · lowest

The description of movement (conveyed, implied, or actual) feels appropriate for this artwork.

4.15SD 0.72
82.0% rated 4 or 5
Q5

The description of materials (oil, canvas, etc.) is plausible given what I can see.

4.35SD 0.79
85.3% rated 4 or 5
Q6

The description of techniques (brushwork, layering, pointillism, etc.) is consistent with the visual evidence.

4.36SD 0.71
89.0% rated 4 or 5
Q7

The emotional expression or mood described by the model feels reasonable for this artwork.

4.29SD 0.77
85.0% rated 4 or 5
Q8

Overall, the set of descriptions gives a coherent and useful interpretation of the artwork.

4.35SD 0.73
88.0% rated 4 or 5
Q9

The descriptions are consistent with what I would expect for the artwork's recorded art style.

4.35SD 0.72
89.7% rated 4 or 5
Q10 · highest

These AI generated descriptions would help me (or a student/curator) better understand or discuss the artwork.

4.38SD 0.76
87.7% rated 4 or 5

Annotators differed in how they used the scale: their grand means were 4.16, 4.62 and 4.20, and a Friedman test indicates a systematic difference in rating levels between them. Across annotator pairs, 85.1% of ratings differ by at most one scale point, while exact agreement is 37.9%.

06

The corpus

23 artists and 34 recorded styles from WikiArt. Since the collection emphasizes historically prominent artists, our analyses describe this corpus rather than art history as a whole.

Artists{{ artistCount }}
{{ a }}
Styles{{ styleCount }}
{{ s }}
07

Interactive figures

Twenty figures from the supplementary material. Each opens in a full screen viewer and loads only when you ask for it, so nothing heavy downloads until you click.

{{ figureCount }}
{{ f.id }} · {{ f.tag }}
{{ f.id }}{{ f.tag }}

{{ f.title }}

{{ f.caption }}

Open figure →

Hover a panel to open it, click to view the figure full screen. Figures are Plotly documents of several megabytes each and load only when opened; on a phone, rotate to landscape.

{{ viewerId }}
{{ viewerTitle }}
{{ viewerTag }}
New tab
{{ viewerCaption }}