DreamLayer Eval Model Leaderboard

Reproducible benchmarks of leading image generation models, run with fixed prompts, seeds, and configs by DreamLayer Eval, the open-source evaluation project from DreamLayer AI.
Compatible with leading image generation APIs and open-source diffusion model workflows
Leaderboard Overview
200 Prompts Benchmarked in 45 Minutes per Model
See how leading image generation models compare across reproducible evaluation metrics such as CLIP Score, FID, and Composition Correctness. DreamLayer Eval automated prompt orchestration, generation, scoring, and result aggregation across models.
‍
Methodology: This benchmark used a prompt set derived from Microsoft COCO and a reference set based on the CIFAR training split. To keep the evaluation controlled and reproducible, the same prompts, seeds, and configs were used across all models. The benchmark was published in September 2025.
CLIP Score
Measures how closely a generated image matches its text prompt.
Rank
Company
Model
CLIP Score
1
Luma Labs
Photon
0.265
2
Black Forest Labs
Flux Pro
0.263
3
OpenAI
Dall-E 3
0.259
4
Google Gemini
Nano Banana
0.258
5
Runway AI
Runway Gen 4
0.2505
6
Ideogram
Ideogram V3
0.2501
7
Stability AI
Stability SD Turbo
0.249
FID Score
Assess how close an AI-generated images are to real images.
Rank
Company
Model
FID Score
1
Ideogram
Ideogram V3
305.60
2
OpenAI
Dall-E 3
306.08
3
Runway AI
Runway Gen 4
317.52
4
Luma Labs
Photon
318.55
5
Black Forest Labs
Flux Pro
318.63
6
Google Gemini
Nano Banana
318.80
7
Stability AI
Stability SD Turbo
321.75
F1 Score
Combines precision and recall to show overall image accuracy.
Rank
Company
Model
F1 Score
1
Luma Labs
Photon
0.463
2
Stability AI
Stability SD Turbo
0.447
3
Runway AI
Runway Gen 4
0.445
4
Black Forest Labs
Flux Pro
0.421
5
Ideogram
Ideogram V3
0.415
6
OpenAI
Dall-E 3
0.380
7
Google Gemini
Nano Banana
0.351
Precision
Measures how many AI-images came out correct when they were compared to the total number of images the AI generated.
Rank
Company
Model
Precision Score
1
Luma Labs
Photon
0.448
2
Stability AI
Stability SD Turbo
0.432
3
Runway AI
Runway Gen 4
0.423
4
Black Forest Labs
Flux Pro
0.406
5
Ideogram
Ideogram V3
0.397
6
OpenAI
Dall-E 3
0.358
7
Google Gemini
Nano Banana
0.339
Recall
Measures how many of the correct images the AI was able to produce out of all the possible correct images it could've generated.
Rank
Company
Model
Recall Score
1
Stability AI
Stability SD Turbo
0.533
2
Luma Labs
Photon
0.532
3
Runway AI
Runway Gen 4
0.522
4
Ideogram
Ideogram V3
0.497
5
Black Forest Labs
Flux Pro
0.495
6
OpenAI
Dall-E 3
0.477
7
Google Gemini
Nano Banana
0.415
CLIP Score
Measures how closely a generated image matches its text prompt.
Rank
Company
CLIP Score
1
Luma Photon
0.265
2
BFL Flux Pro
0.263
3
OpenAI Dall-E 3
0.259
4
Google Nano Banana
0.258
5
Runway Gen 4
0.2505
6
Ideogram V3
0.2501
7
Stability SD Turbo
0.249
FID Score
Assess how close an AI-generated images are to real images.
Rank
Company
FID Score
1
Ideogram V3
305.60
2
OpenAI Dall-E 3
306.08
3
Runway Gen 4
317.52
4
Luma Photon
318.55
5
BFL Flux Pro
318.63
6
Google Nano Banana
318.80
7
Stability SD Turbo
321.75
F1 Score
Combines precision and recall to show overall image accuracy.
Rank
Company
F1 Score
1
Luma Photon
0.463
2
Stability SD Turbo
0.447
3
Runway Gen 4
0.445
4
BFL Flux Pro
0.421
5
Ideogram V3
0.415
6
OpenAI Dall-E 3
0.380
7
Google Nano Banana
0.351
Precision
Measures how many AI-images came out correct when they were compared to the total number of images the AI generated.
Rank
Company
Precision Score
1
Luma Photon
0.448
2
Stability SD Turbo
0.432
3
Runway Gen 4
0.423
4
BFL Flux Pro
0.406
5
Ideogram V3
0.397
6
OpenAI Dall-E 3
0.358
7
Google Nano Banana
0.339
Recall
Measures how many of the correct images the AI was able to produce out of all the possible correct images it could've generated.
Rank
Company
Recall Score
1
Stability SD Turbo
0.533
2
Luma Photon
0.532
3
Runway Gen 4
0.522
4
Ideogram V3
0.497
5
BFL Flux Pro
0.495
6
OpenAI Dall-E 3
0.477
7
Google Nano Banana
0.415

FAQ

DreamLayer AI and DreamLayer Eval

What is DreamLayer AI?

DreamLayer AI is an image generation and editing platform. You describe the image you need, add a reference when identity or style should carry forward, and refine the result over several turns, in the browser or through an API, CLI, and MCP server. This leaderboard comes from DreamLayer Eval, its separate open-source benchmarking project.

Which AI models does DreamLayer use?

DreamLayer AI routes each request to one of several image models depending on the task. The models on this page were benchmarked by DreamLayer Eval in September 2025; the product's model choices are maintained separately and are not tied to this snapshot.

How is DreamLayer different from a single AI image generator?

A single-model tool uses one model for every request. DreamLayer AI chooses among several image models per task, carries a reference identity forward across edits, and offers background removal and upscaling in the same conversation.

Can DreamLayer edit just one part of an image?

Yes. Upload the image as a reference and describe the change you want. DreamLayer AI applies the edit and keeps the rest of the image, and you can refine the result over further turns.

Do I need my own API keys or a GPU to use DreamLayer?

No. DreamLayer AI is a hosted product. It runs in the browser and through its API, CLI, and MCP server, and it manages model access for you. DreamLayer Eval, the open-source benchmarking project, is the part you run locally if you want to reproduce benchmarks.

How much does DreamLayer cost?

One credit per finished image, across the browser, API, CLI, and MCP server. Credits do not expire, and an unfinished request restores the reserved credit. Details are on the pricing page.

Open-Source Benchmarking

What is DreamLayer's open-source benchmarking tool?

DreamLayer Eval is benchmarking infrastructure for image and video diffusion models. It automates prompts, seeds, configs, metric scoring, and reproducible run logging so researchers can compare models consistently. The code is available on GitHub.

What metrics does the benchmarking tool support?

Built-in evaluation metrics include CLIP Score, FID, precision, recall, F1, LPIPS, SSIM, PSNR, and temporal consistency for video, all logged automatically with the prompts, seeds, and configs that produced them.

Are DreamLayer benchmark results reproducible?

Yes. Every run is logged with its prompts, seeds, configs, outputs, and metric scores, and can be exported as CSV, JSON, or a complete benchmark bundle, so results stay traceable and repeatable.

How do the benchmarks relate to the DreamLayer product?

DreamLayer Eval measures how models perform under fixed prompts, seeds, and configs, and the leaderboard on this page is its September 2025 result. DreamLayer AI is the separate hosted product. It does not depend on DreamLayer Eval, and its model choices are maintained on their own.