Gpt is really good in vision stuff, or at least their MoE seems to be really cohesive. From my experience Claude models can be really good at language but the moment they need to look at a picture and decide why the design is not good what parts need improvement it degrades a lot. My easiest benchmark is giving them a screenshot of a feature in my app and tell it "identify non-normative UI blocks and improve readability and consistency". Sol does a great job at re-structuring the page into composable units that build upon each other and the general looks and feels of the app. Claude tends to over-focus one one part while completely forgetting about the rest or the cohesion as a whole.
If you're doing any kind of inference that is multi-modal and non-factual, opinions and biases will affect any kind of assessment of a visual that you provide to a model.
For example, a UI / UX professional being asked to appraise a website screenshot may determine that the image in question has "desirable" traits which are inherently not deterministically measurable. Such as, if the interface elements have strong information hierarchy, or if they are deemed to be "fashionable" with current UI trends.
As a design system engineer I usually have to fight against the taste of the designers. (And I consider it natural.)
But, if you have a proper well documented design system and you tell the LLM to use the DS and to avoid styling hacks they can generally do it. Even the dumber ones than Sol 5.6.
Of course only if the design is achievable in the design system.
For 2 weeks I've been trying to get Codex to "outpaint" a wonderful image it generated as placeholder art for a level background.
After I increased the game's resolution, I asked it to increase the image's size while keeping the same scale and existing content, and gosh, it constantly keeps getting something wrong no matter what I tell it, even on Sol Max with the $100 Pro subscription.
An average pixel-artist could have recreated the image and more within 2-3 days.
[delayed]
Anecdotal, opinion:
Gpt is really good in vision stuff, or at least their MoE seems to be really cohesive. From my experience Claude models can be really good at language but the moment they need to look at a picture and decide why the design is not good what parts need improvement it degrades a lot. My easiest benchmark is giving them a screenshot of a feature in my app and tell it "identify non-normative UI blocks and improve readability and consistency". Sol does a great job at re-structuring the page into composable units that build upon each other and the general looks and feels of the app. Claude tends to over-focus one one part while completely forgetting about the rest or the cohesion as a whole.
Assessing the subjective quality of a thing is in my experience one of the worst ways to use any LLM.
anthropic frontend-design skill does a great job with it.
What is a "non-normative UI block"?
areas that look weird
Penny sample shown looks like failed EXIF orientation registered by the model/harness. The coins are correctly marked, it's rotated 90 degrees.
In the third vision bench result, Sol is 100% correct but the expected has 1 error. Seems like an oversight.
In the next bench, Sol looks like it’s correct again but the bboxes are rotated 90 degrees for some reason.
My anecdotal evidence says its still as blind as any other model, it has no taste, no attention to any sort of detail.
How can a vision model have taste?
If you're doing any kind of inference that is multi-modal and non-factual, opinions and biases will affect any kind of assessment of a visual that you provide to a model.
For example, a UI / UX professional being asked to appraise a website screenshot may determine that the image in question has "desirable" traits which are inherently not deterministically measurable. Such as, if the interface elements have strong information hierarchy, or if they are deemed to be "fashionable" with current UI trends.
> if the interface elements have strong information hierarchy
...but that's an example of a UX/usability matter that can be assessed objectively and non-subjectively.
Replace taste with consistent if that helps you. Can it follow a design system...
As a design system engineer I usually have to fight against the taste of the designers. (And I consider it natural.)
But, if you have a proper well documented design system and you tell the LLM to use the DS and to avoid styling hacks they can generally do it. Even the dumber ones than Sol 5.6.
Of course only if the design is achievable in the design system.
This is not my experience at all.
So, formulaic output…the opposite of taste
Not really. Compliance with the letter of the law doesn't mean the intent is complied with.
For 2 weeks I've been trying to get Codex to "outpaint" a wonderful image it generated as placeholder art for a level background.
After I increased the game's resolution, I asked it to increase the image's size while keeping the same scale and existing content, and gosh, it constantly keeps getting something wrong no matter what I tell it, even on Sol Max with the $100 Pro subscription.
An average pixel-artist could have recreated the image and more within 2-3 days.
> it constantly keeps getting something wrong no matter what I tell it
This 100%
did you try segmenting it first?