Natural forms remain prominent, while geometric and irregular forms become more frequent in later samples; reduced scales also become more visible over time.
Hi, and welcome
You are in the right place. This page is the companion to the paper: the abstract, the method, what we found, and every interactive figure from the supplementary material, ready to explore.
Multimodal language models can increasingly describe both the visual and expressive qualities of artworks, raising questions about how aesthetic interpretation changes when part of the process is delegated to AI. We present CognArtive, a framework that applies a structured art analysis protocol to more than 15,000 artworks from 23 artists and 34 styles. GPT vision generates artwork level interpretations across dimensions including form, color, light, movement, material, technique, temporal cues, figures, and emotional expression, while GPT and Gemini support synthesis and collection scale analysis. We evaluate how these model generated descriptions align semantically with established art style descriptions using four embedding models and find substantial variation across aesthetic dimensions. Beyond large scale analysis, CognArtive provides a case study of interpretive agency in creative AI, showing how human selected criteria, model representations, canonical categories, and corpus composition jointly shape computational interpretations of art.
The pipeline
Human selected questions define the analytical dimensions, GPT 4V interprets the visual evidence, and GPT 4 with Gemini 2.0 synthesize the outputs for collection scale analysis, semantic evaluation, and interactive visualization.
Artwork corpus
More than 15,000 digitized artworks from WikiArt: 23 artists, 34 recorded styles, roughly the fifteenth to the twenty first century.
Eight questions
A human defined framework derived from formal art analysis. Seven questions follow Hodge (2024); the eighth identifies the presence and type of figures.
GPT 4V interpretation
Each artwork image and the eight questions go to GPT 4V, which produces descriptions of the corresponding aesthetic dimensions.
Synthesis
Responses are cleaned, then processed with GPT 4 and Gemini 2.0 into structured representations for qualitative reading and quantitative aggregation.
Three outputs
Collection scale analysis across artists, styles and periods; semantic comparison with art style descriptions; and the interactive figures on this page.
Eight questions, one artwork at a time
Every artwork is put to the same eight questions. Hover a spoke to hold it, and the panel shows what the model returns for that dimension, what it found across the collection, and how well it agrees with art historical language.
{{ qDesc }}
{{ qFinding }}
Twelve tones
Twelve tonal categories recur across the six periods. Hover a band for its register. Swatches are illustrative; the measured distributions are in Figure S5.
What we found
The patterns below describe model generated annotations within the sampled corpus. Each card opens the interactive figure behind it.
Color distributions show broad temporal variation, with monochromatic tones stronger in earlier samples and muted tones more visible later.
CognArtive identifies more than twenty emotional themes, with positive and neutral interpretations occurring more frequently than negative ones.
Light and contrast are frequently associated with highlights, depth, and texture, while contrast, shadow, and chiaroscuro are common visual effects. Diffused and soft lighting appear frequently, and emphasis or directing attention are common inferred purposes across major styles.
Conveyed and implied movement occur more frequently than literal movement.
Oil and canvas dominate material descriptions, while other media show distinct temporal patterns. Blending and layering are among the most frequent techniques, alongside variations in brushwork, line, crosshatching, scraping, and pointillism.
Morning and afternoon are common temporal cues, while human figures dominate the represented subjects.
Three annotators rated 100 randomly sampled artworks across ten criteria: a mean of 4.33 out of 5, with 87.4% of 3,000 ratings scored 4 or 5.
See the full ratings →Semantic alignment
Median cosine similarity between CognArtive descriptions and established art style descriptions, across four embedding models. Techniques show the strongest aggregate alignment, followed by Movement and Material.
{{ readNote }}
Aggregate reports the median after aggregation across embedding models. BGE M3 and NV Embed v2 return consistently higher similarities than Sentence BERT and OpenAI for every dimension, so the choice of embedding space shifts the absolute values more than the ranking between dimensions.
Ten questions, three annotators
Each annotator rated all 100 randomly sampled artworks on a five point Likert scale, producing 3,000 individual ratings.
The description of form and scale accurately matches what I see in the artwork.
The description of colors accurately matches the colors I see in the artwork.
The description of light and contrast accurately reflects the lighting and contrast in the artwork.
The description of movement (conveyed, implied, or actual) feels appropriate for this artwork.
The description of materials (oil, canvas, etc.) is plausible given what I can see.
The description of techniques (brushwork, layering, pointillism, etc.) is consistent with the visual evidence.
The emotional expression or mood described by the model feels reasonable for this artwork.
Overall, the set of descriptions gives a coherent and useful interpretation of the artwork.
The descriptions are consistent with what I would expect for the artwork's recorded art style.
These AI generated descriptions would help me (or a student/curator) better understand or discuss the artwork.
Annotators differed in how they used the scale: their grand means were 4.16, 4.62 and 4.20, and a Friedman test indicates a systematic difference in rating levels between them. Across annotator pairs, 85.1% of ratings differ by at most one scale point, while exact agreement is 37.9%.
The corpus
23 artists and 34 recorded styles from WikiArt. Since the collection emphasizes historically prominent artists, our analyses describe this corpus rather than art history as a whole.
Interactive figures
Twenty figures from the supplementary material. Each opens in a full screen viewer and loads only when you ask for it, so nothing heavy downloads until you click.
Hover a panel to open it, click to view the figure full screen. Figures are Plotly documents of several megabytes each and load only when opened; on a phone, rotate to landscape.