AI Image Generation for the Web: Quality, Licensing and Performance
AI image generation for real websites: the quality tells that give it away, what generated images licensing actually gives you, and how to cut file size 30%.
A client sent us a hero image last month that they’d generated themselves and were very proud of. It looked great in Figma. It was a 1024×1024 PNG weighing 2.1MB, upscaled to 2560px wide in Preview, and when we dropped it into the staging build the Largest Contentful Paint on a throttled 4G profile went from 1.6s to 7.4s. The image also contained a shop sign reading “COFFEE BREWBRERY” and a man holding a cup with four knuckles on his thumb.
Both problems are fixable. Neither is obvious if your only experience of generated imagery is admiring it at 50% zoom on a Retina display. The third problem, the one nobody noticed until legal asked, was that nobody could say with confidence who owned the thing.
This is the version of the topic written after shipping generated assets on real client sites, getting it wrong twice, and building a pipeline that now handles it without drama.
- Purely machine-generated images generally are not protectable by copyright in the US, which means you can use them but you often cannot stop a competitor from using the identical output. For logos and brand marks, that’s disqualifying.
- Diffusion output carries high-frequency noise that wrecks modern codecs. A light median filter or 0.3px blur before AVIF encoding typically cuts file size by 20 to 35% with no visible difference.
- Never serve the generator’s native PNG. A 1024px PNG at around 2MB becomes an 80 to 140KB AVIF at quality 50, and that’s the file your LCP cares about.
- Indemnification is the real differentiator between providers, not image quality. Adobe Firefly, Getty and Shutterstock offer it on commercial tiers; most open-weight models offer nothing, and Flux.1 [dev] is explicitly non-commercial.
- The EU AI Act’s transparency duties for synthetic media bite from August 2026, so machine-readable marking (C2PA Content Credentials, IPTC
trainedAlgorithmicMedia) stops being optional for EU-facing sites.
What AI image generation is actually good at
The useful framing isn’t “can it replace a photographer”. It’s “which asset slots on this site tolerate an image that nobody will look at closely”. That list is longer than purists admit and shorter than vendors claim.
Things we now ship generated routinely: abstract background textures and gradients, blog post header art for topics with no natural photo, 404 and empty-state illustrations, pattern fills, blurred bokeh plates behind testimonial sections, product mood boards during the design phase, and icon sets in a consistent style. Empty states are a particularly good fit. They’re low-traffic, need to feel light, and commissioning illustration for them never survives a budget review. We wrote about why those screens deserve attention at all in designing for failure.
Things that keep biting us: anything with legible text (Ideogram and gpt-image-1 are far better than they were, still not reliable at small sizes), human hands at close range, specific real products, architectural interiors with consistent reflections, anything that must match a photographed set, and food. Food looks plasticky in a way that reads as fake to people who can’t articulate why.
The hard rule we use: if a user might plausibly zoom in, don’t generate it. If the image is making a factual claim about your product, don’t generate it. Showing a generated screenshot of software that doesn’t look like your software isn’t a style choice, it’s a misrepresentation, and the support tickets arrive within a week.

The quality tells that get you caught
Visitors rarely say “that’s AI”. They say the site feels generic. The tells are mostly structural rather than content-level:
- Uniform shallow depth of field. Diffusion models love f/1.4 bokeh. Three hero images in a row with the same creamy background and the page reads as stock, synthetic or both.
- Symmetric, centred composition. Models drift toward the subject dead-centre with equal negative space. Real photography has off-balance crops. Ask for them explicitly, or crop hard in post.
- Light with no source. Rim lighting on all sides, shadows pointing in two directions, specular highlights that don’t match the window. This is the single most common reason a composite looks wrong.
- Over-detailed everywhere. Real lenses have a focal plane. Generated images often render background foliage at the same sharpness as the foreground, which triggers an uncanny response and, conveniently, also destroys your compression ratio.
- Texture mush at 100%. Fabric weaves, brick, hair at the edges. Fine at 1x, smeared at 2x. If you’re generating at 1024 and serving at 2x for a 1280px-wide hero, you’re showing the mush.
Generate at or above your largest served size. Most current APIs will do 1536px on the long edge; Flux and SDXL-class models handle 1920×1080 natively. Upscaling with Real-ESRGAN or Topaz is acceptable for texture plates and bad for faces, because upscalers hallucinate detail that doesn’t match the original structure.
Generated images licensing: what you actually own
This is where most teams are exposed, and it’s not about training data lawsuits. It’s about your own rights in the output.
In the US, the Copyright Office’s position (consistent through Zarya of the Dawn, the Thaler litigation and its 2025 report on copyrightability) is that material generated by a machine in response to a prompt lacks human authorship and isn’t registrable. Prompting, however elaborate, isn’t authorship. Human-authored arrangement, editing and selection can attract protection in those contributions, not in the underlying generated pixels.
Practical consequence: you can put a generated image on your marketing site all day. You generally cannot register it, cannot assign it with warranties you can actually honour, and cannot stop a competitor who runs a near-identical prompt and gets a near-identical result. For a hero background, nobody cares. For a logo, a character mascot, a packaging illustration or anything you’d ever want to enforce, generated output is the wrong choice. We’ve had to re-commission a mascot for exactly this reason, after it was already on the van.
Provider terms, ranked by how much they protect you
- Adobe Firefly, Getty’s generative tool, Shutterstock AI: trained on licensed or owned libraries, and the commercial or enterprise tiers carry indemnification. If an third-party claim lands, somebody else’s lawyers answer the phone. This is why enterprise clients keep choosing Firefly even when the output is blander.
- OpenAI, Google, Midjourney paid tiers: terms assign you whatever rights they can in the output and let you use it commercially. Midjourney specifically requires a Pro-level plan if your company turns over more than $1M a year, and free-tier output falls under a non-commercial Creative Commons grant. Also worth knowing: unless you’re on Stealth mode, your generations are public.
- Open weights: read the licence, every time. Flux.1 [schnell] is Apache 2.0 and fine commercially. Flux.1 [dev] is a non-commercial licence, and we have seen it in production on sites whose owners had no idea. Stable Diffusion 3.5 has its own community licence with a revenue threshold.
Keep a CSV: image filename, model, date, prompt, provider plan, licence note. When a client’s acquirer runs IP diligence in two years, that file is the difference between a tidy answer and a frantic afternoon. Treat it exactly like the asset provenance records you keep for any licensed photography.
Provenance, disclosure and the 2026 rules
Two things changed the calculus. C2PA Content Credentials moved from demo to default in Adobe tooling, OpenAI’s image outputs and several camera bodies. The IPTC vocabulary now gives you a precise value for “made by a model”: digitalSourceType: trainedAlgorithmicMedia. And the EU AI Act’s Article 50 transparency obligations for synthetic content apply from 2 August 2026, requiring machine-readable marking of AI-generated image, audio and video content. If you serve EU users, this is now a build requirement, not a position.
The practical problem is that almost every image pipeline strips metadata. sharp drops it by default. CDNs resize and re-encode. WordPress regenerates thumbnails. So “we keep the Content Credentials” quietly becomes false somewhere between the designer’s export and the browser.
const sharp = require('sharp');
await sharp('src/hero-raw.png')
// keepMetadata preserves XMP/IPTC, including C2PA assertions and
// digitalSourceType, which sharp otherwise discards on re-encode
.keepMetadata()
.median(1) // kills diffusion micro-noise, big AVIF win
.resize({ width: 1920 })
.avif({ quality: 50, effort: 5, chromaSubsampling: '4:2:0' })
.toFile('dist/hero-1920.avif');
On the user-facing side, we don’t label every background texture. We do disclose anywhere the image could be mistaken for documentary evidence: team photos, case study imagery, product shots, anything with a human face. A single line in the image caption or page footer is enough, and it costs you nothing in trust. Hiding it costs a lot when someone notices.
Why generated images are born heavy
Here’s the part the listicles skip. Diffusion output is not just large, it’s expensively large. The denoising process leaves a fine, non-structured grain across the whole frame. Modern codecs work by predicting blocks from neighbours, and that grain is unpredictable by construction. Same visual complexity, worse compression.
Measured on a batch of 40 generated 1536px hero plates on a recent build:
- Source PNG from the API: 1.4MB to 2.6MB each.
- Straight AVIF, quality 50: average 212KB.
- Median filter radius 1, then AVIF quality 50: average 148KB. That’s a 30% reduction with no perceptible difference at 1x or 2x.
- WebP quality 78 for the fallback: average 231KB.
A 0.3px Gaussian blur gets you most of the same benefit and is gentler on fine edges like typography or wires. Try both on your actual set; the winner depends on the model. Flux output is grainier than gpt-image-1 in our testing, and benefits more.
Then the standard discipline applies, and generated images make it more important rather than less, because the temptation is to use a big decorative hero on every page:
<picture>
<source type="image/avif"
srcset="/img/hero-960.avif 960w, /img/hero-1440.avif 1440w, /img/hero-1920.avif 1920w"
sizes="(max-width: 991px) 100vw, 1140px">
<source type="image/webp"
srcset="/img/hero-960.webp 960w, /img/hero-1440.webp 1440w, /img/hero-1920.webp 1920w"
sizes="(max-width: 991px) 100vw, 1140px">
<img src="/img/hero-1440.jpg" alt="Abstract layered paper texture in warm neutrals"
width="1920" height="1080"
fetchpriority="high" decoding="async">
</picture>
No loading="lazy" on the LCP element. Ever. We still find it there on about one in three sites we audit, usually because a plugin added it globally. Explicit width and height to hold layout. And if the generated image is purely decorative behind text, consider whether it should be a CSS background that only loads above a breakpoint, or whether a CSS gradient would do the same job for 0KB. Often it would. The Bootstrap 5 demos in Canvas lean on gradient and SVG pattern backgrounds for exactly this reason: they scale to any viewport and never show up in a waterfall.
A pipeline that survives a real build
Generating in a browser tab and dragging files into /img works for one image and collapses at twenty. What we run now:
- Generate to a staging folder with a manifest. One JSON record per image: prompt, seed, model, version, provider, plan, licence, date, target slot. Written by the generation script, not by hand. If you’re calling an image API from code, the same discipline about validated, schema-shaped responses applies as with text models: see structured output from LLMs.
- Human review gate. A contact sheet at 100% crop on faces, hands and any text region. Rejects go back with a note. This takes four minutes for twenty images and catches the BREWBRERY sign.
- Process, don’t trust. Denoise, resize to the actual breakpoints, encode AVIF plus WebP plus a JPEG fallback, preserve metadata, write dimensions into the manifest.
- Alt text written by a human, drafted by a model. A vision model will describe the image accurately, and that description is almost never the right alt text, because alt text depends on the image’s function on the page. Use the draft as a starting point. Decorative plates get
alt="". - Budget check in CI. Fail the build if any image over 200KB lands in the hero slot, or if total page image weight crosses your budget. A Lighthouse CI assertion or a simple file-size script both work.
If this is your first AI-adjacent thing in a production pipeline, the generic shape of the problem (review gates, fallbacks, measuring before and after) is the same one we covered in shipping an AI feature users trust.
Keeping a set consistent
One good image is easy. Nine images that look like they came from the same photographer is the actual job, and it’s where most generated image sets on a website fall apart. Visitors can’t name the inconsistency but they register it as cheapness.
What works, in order of effectiveness:
- Style references over prompt text. Midjourney’s
--sref, Flux Redux, or an IP-Adapter conditioning image beat any amount of adjective stacking. Pick one reference frame, use it for the whole set. - Lock the seed, vary the subject. Keeps lighting and grain coherent across a grid.
- Train a small LoRA if the set is big enough to justify it. Fifteen to thirty images of your real product or real office, a couple of hours of training on a rented A100, and you get output that actually matches your brand instead of matching the internet’s average. This is the step that turns generated imagery from a shortcut into a genuine capability.
- Finish in Photoshop or Affinity. Colour-grade the whole set with one LUT. Add matched grain after denoising. Crop to a consistent aspect and a consistent subject position. Twenty minutes here does more for perceived quality than another fifty generations.
Accept the limit: if the client has a photographed product catalogue, generated lifestyle shots around it will never quite match. Mixing the two in one grid looks worse than either alone. Pick a lane per page.
Frequently Asked Questions
Does Google penalise AI images on a website?
No. Google’s spam policies target content produced at scale to manipulate rankings, not the tool used to make an image. Generated imagery that illustrates the page usefully is treated like any other image, and the things that actually affect ranking (file size, LCP, alt text, structured data, surrounding context) are unchanged. Google does read C2PA Content Credentials and can surface them in “About this image”, which is an argument for preserving metadata rather than stripping it.
Can I use a generated image as my company logo?
You can, but you shouldn’t. Purely generated output almost certainly isn’t copyrightable in the US, so you’d have no copyright to enforce if a competitor used something near-identical. You may still be able to get trademark protection through use, since trademark rights come from commerce rather than authorship, but you’re building a brand asset on a weaker legal footing than a commissioned mark. Pay a designer for the logo and generate the blog headers.
What’s the best format for generated hero images in 2026?
AVIF first, WebP as the second source, JPEG as the img fallback. AVIF support is now effectively universal across current Chrome, Firefox, Safari and Edge, but the JPEG fallback costs you nothing in a picture element and covers old in-app browsers. Encode AVIF at quality 45 to 55 with effort 4 to 6; higher effort is slower to encode and barely smaller.
Do I need to disclose that images are AI-generated?
If you serve EU users, machine-readable marking of synthetic content is required under the AI Act’s transparency obligations from August 2026, and the safe implementation is to keep C2PA credentials and IPTC trainedAlgorithmicMedia intact through your build. Visible disclosure isn’t universally mandated, but we recommend it anywhere the image could be read as documentary: people, premises, product shots, case studies. Decorative textures don’t need a label.
Is generating images cheaper than stock photography?
Per image, yes, usually by an order of magnitude at a few cents per generation versus $10 to $80 for a decent stock licence. Per usable set, it’s closer than you’d think, because you’ll generate fifteen to forty candidates per keeper and then spend real designer time grading and cropping. Where generation wins outright is on things stock libraries handle badly: abstract backgrounds, specific illustration styles and imagery for narrow technical subjects.
So what should you actually do
Pick your slots deliberately. Backgrounds, textures, blog headers, empty states: generate, and build the denoise-plus-AVIF step into your pipeline before you generate anything, because retrofitting it across 60 images is miserable. Logos, product shots, team photos, anything a court or a customer might rely on: don’t.
Then do the boring bit. Start the manifest today, even if it’s a spreadsheet with five rows, and record the model and licence for every generated asset that ships. Nobody has ever regretted having that file. Plenty of teams have regretted not having it.


