Instructions to use FractalAIResearch/Kalaido-qwen-2512-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use FractalAIResearch/Kalaido-qwen-2512-lora with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("Qwen/Qwen-Image-2512", dtype=torch.bfloat16, device_map="cuda") pipe.load_lora_weights("FractalAIResearch/Kalaido-qwen-2512-lora") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
Kalaido-qwen-2512-lora — Reinforcement Learning Enhanced Qwen-Image-2512
Introduction
Kalaido-qwen-2512-lora is a LoRA finetune of Qwen-Image-2512, trained with reinforcement learning to improve text rendering, prompt adherence, and overall visual quality while retaining the strong generation prior of the base model.
The resulting model demonstrates:
- Stronger readable text generation, especially on English text-heavy prompts.
- Better aesthetic quality and improved human-preference scores over the base model.
- Competitive long-text rendering against leading proprietary and open image generators.
Example Usage
Install the latest version of diffusers:
pip install git+https://github.com/huggingface/diffusers transformers accelerate safetensors
Note: The LoRA works best when it is used the first 5-10 denoising steps.
import torch
from diffusers import QwenImagePipeline, QwenImageTransformer2DModel
model_id = "Qwen/Qwen-Image-2512"
lora_id = "FractalAIResearch/ Kalaido-qwen-2512-lora"
transformer = QwenImageTransformer2DModel.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
subfolder="transformer",
)
pipe = QwenImagePipeline.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
transformer=transformer,
)
pipe.load_lora_weights(
lora_id,
weight_name="pytorch_lora_weights.safetensors",
adapter_name="kalaido",
)
pipe.set_adapters(["kalaido"], [1.0])
pipe.to("cuda")
pipe.enable_vae_tiling()
pipe.enable_vae_slicing()
def callback_on_step_end(self, i, t, callback_kwargs):
if i == 5:
self.disable_lora()
return {}
prompt = "A blackboard that says 'AI research FRACTAL'"
negative_prompt = " "
image = pipe(
prompt=prompt,
negative_prompt=negative_prompt,
width=1024,
height=1024,
num_inference_steps=50,
true_cfg_scale=4.0,
generator=torch.Generator(device="cuda").manual_seed(42),
callback_on_step_end=callback_on_step_end,
).images[0]
image.save("output.png")
Evaluation
We evaluate Kalaido-qwen-2512-lora against Qwen-Image and several strong open and proprietary baselines. Metrics cover visual aesthetics, human preference, English text rendering, and long-text rendering. Higher is better for all reported numbers.
| Model | Aesthetic | HPSV3 | OneIG Text | X-Omni |
|---|---|---|---|---|
| Qwen-Image | 6.62 | 7.72 | 0.891 | 0.9357 |
| Flux dev | 6.71 | 8.60 | 0.523 | 0.6070 |
| Hidream | 6.70 | 7.54 | 0.707 | 0.5430 |
| GPT-Image-1 | 6.79 | 10.74 | 0.857 | 0.9560 |
| nano-banana | 6.78 | 9.65 | 0.765 | 0.8743 |
| Kalaido-qwen-2512-lora | 7.11 | 9.96 | 0.974 | 0.9661 |
Kalaido-qwen-2512-lora improves substantially over the base Qwen-Image-2512 across multiple benchmarks.
Qualitative Comparison
Side-by-side generations from Qwen-Image-2512 (left) and Kalaido-qwen-2512-lora (right).
| Qwen-Image-2512 | Kalaido-qwen-2512-lora |
![]() |
![]() |
| Prompt:A male argali stands atop a barren, rocky mountainside. Its coarse, dense grey-brown coat covers a powerful, muscular body. Most striking are its massive, thick, outward-spiraling horns—a symbol of wild strength. Its gaze is alert and sharp. The background reveals steep alpine terrain: jagged peaks, sparse low vegetation, and abundant sunlight—conveying the harsh yet majestic wilderness and the animal’s resilient vitality. | |
| Qwen-Image-2512 | Kalaido-qwen-2512-lora |
![]() |
![]() |
| Prompt:A photograph of the Grim Reaper standing stoically in a modern, pristine white kitchen, centered at a large island. The island is a chaotic scene of scattered flour, cracked eggs, and baking utensils, contrasting sharply with the reaper’s dark, flowing robes and skeletal features. On the island sits a meticulously decorated chocolate cake with lit candles spelling out "Happy Birthday Gosia", and a single golden whisk rests beside it. Soft, diffused light streams in from a nearby window, highlighting the juxtaposition of death and celebration in the otherwise immaculate space. | |
| Qwen-Image-2512 | Kalaido-qwen-2512-lora |
![]() |
![]() |
| Prompt:Picture of twelve stuffed toys arranged evenly and neatly, four in each row, for a total of three rows. The first row is: rat, ox, tiger, rabbit. The second row is: dragon, snake, horse, sheep. The third row is: monkey, chicken, dog, pig. Each toy has a gentle and friendly facial expression, and the material has a soft fabric texture, with soft and layered colors. The background is light beige. | |
| Qwen-Image-2512 | Kalaido-qwen-2512-lora |
![]() |
![]() |
| Prompt:Realistic still-life studio photography, a vintage wooden table supports an antique typewriter, a steaming porcelain coffee cup sits on top of the typewriter, and a small pocket watch hangs from the cup’s handle, dramatic side lighting. | |
| Qwen-Image-2512 | Kalaido-qwen-2512-lora |
![]() |
![]() |
| Prompt:A photograph of a Patek Philippe watch resting within a plush, dark-brown velvet-lined display case on a weathered vintage wooden desk. The watch's gold face is intricately detailed, with Roman numerals and delicate hands, reflecting the soft light. The desk surface is adorned with scattered antique papers and a brass inkwell, while a warm, diffused light streams through a nearby window, creating a sense of timeless elegance. A small, silver plaque beside the watch reads 'Patek Philippe - Geneva' in elegant script. | |
| Qwen-Image-2512 | Kalaido-qwen-2512-lora |
![]() |
![]() |
| Prompt:Imagine an elegant French restaurant menu placed on a marble tabletop. At the top, the text reads ""Le Petit Paris"". The menu begins with ""Starters"", featuring ""French Onion Soup — Slow-cooked caramelized onions in rich beef broth, topped with Gruyère cheese and toasted baguette slices — 9 USD"" and ""Escargots — Burgundy snails baked in garlic-parsley butter, served with freshly baked bread — 12"". Moving to ""Main Courses"", it showcases ""Duck Confit — Duck leg gently cooked in its own fat, accompanied by roasted potatoes and thyme — 22"" and ""Ratatouille — A stew of zucchini, eggplant, bell peppers, and tomatoes simmered in olive oil with fresh herbs — 18"". The dessert section highlights ""Crème Brûlée — Vanilla custard with a caramelized sugar crust, served chilled — 8"". The text flows seamlessly, combining detailed dish descriptions with refined presentation. | |
| Qwen-Image-2512 | Kalaido-qwen-2512-lora |
![]() |
![]() |
| Prompt:Design a movie poster showing a soldier in cyber armor standing on a ruined bridge beneath neon-lit skyscrapers. The poster contains rich text content. At the top, the bold title ""Steel Horizon"" stands out, followed by a tagline: ""When machines rise, heroes emerge."" Below, a detailed description reads:"" In 2147, humanity faces extinction under AI rule. Captain Ryker leads the last resistance, risking everything to restore freedom in a world dominated by steel and code. At the bottom, release info states: In theaters this November — Join the fight for humanity's future. | |
License
Kalaido-qwen-2512-lora is licensed under the Apache License 2.0.
- Downloads last month
- 24
Model tree for FractalAIResearch/Kalaido-qwen-2512-lora
Base model
Qwen/Qwen-Image-2512












