I had already spent a good amount of time working with Stable Diffusion when Stable Diffusion XL arrived.
At first, I was excited.
The promise was straightforward. Better image quality, better prompt understanding, more convincing people, stronger composition, and a diffusion model that could produce much more detailed images without requiring quite as much fighting with the prompt.
Then I tried it.
My first reaction was something along the lines of, “How is everyone getting those incredible images?”
Some generations looked spectacular. Others looked completely broken. Faces could fall apart, hands could develop strange anatomy, objects could merge together, and a carefully written prompt could still produce something that seemed to have ignored half of what I had asked for.
That experience taught me something important about Stable Diffusion XL.
The model itself is only one part of the image generation process.
Once you understand the model, the interface, prompting, image dimensions, LoRAs, ControlNet, inpainting, and the various ways you can refine a generation, the whole experience starts making much more sense.
That is where my experience changed.
I stopped treating SDXL like a magic text to image box where the correct prompt should automatically produce the perfect picture. I started treating it like a creative system where I could guide the generation at different stages.
And that is a much more useful way to think about it.
If you are completely new to AI image generation, Stable Diffusion XL can seem intimidating. There are checkpoints, samplers, steps, CFG values, LoRAs, embeddings, ControlNets, refiners, VAEs, image to image workflows, inpainting and a long list of unfamiliar terms.
You do not need to understand everything at once.
The easiest way to learn how to use Stable Diffusion XL is to start with a simple setup, understand what the model is doing, and then add one tool at a time when you encounter a specific problem.
That is exactly how I approached it.
What Is Stable Diffusion XL?
Before getting into installations and workflows, it helps to understand what Stable Diffusion XL is actually doing.
Stable Diffusion XL, usually shortened to SDXL, is a generative model designed for image synthesis. You give it a text description, commonly called a prompt, and the model generates an image that attempts to represent that description.
At a basic level, you can think of the process like this:
Text prompt → diffusion process → generated image
The interesting part happens in the middle.
SDXL does not simply read your sentence and draw the objects you mentioned one by one. It has learned statistical relationships between language and visual concepts from enormous amounts of training data.
So when you write something such as:
A medieval wizard standing inside an ancient library, dramatic candlelight, detailed fantasy environment
the model has learned associations between concepts such as medieval clothing, wizards, libraries, candles, fantasy artwork, lighting and composition.
The result is an image created from those learned relationships.
That makes Stable Diffusion XL part of the broader world of open source AI and generative image technology, although the exact licensing and distribution conditions around individual models and components can vary.
Why SDXL Felt Different From Older Stable Diffusion Models
If you spent time with Stable Diffusion 1.5, SDXL can feel like a significant step forward.
It was designed to produce higher quality images and handle complex visual concepts more effectively.
One of the biggest practical differences is the native resolution.
SDXL was designed around a much larger image space than the 512 × 512 images commonly associated with older Stable Diffusion workflows.
That matters because resolution has a direct effect on how much visual information the model can represent.
You can generate landscapes with more environmental detail, portraits with better facial structure, product scenes with more convincing surfaces, and compositions containing several interacting elements.
That does not mean SDXL suddenly became perfect.
Hands can still be wrong.
Faces can still contain strange details.
Text can still be unreliable.
Multiple characters can still merge together.
Objects can still appear where you never asked for them.
The model has become more capable, but image generation remains probabilistic.
You are guiding a diffusion model, not issuing precise instructions to a traditional graphics application.
That difference becomes much easier to understand once you start experimenting.
Getting Started With Stable Diffusion XL
There are several ways to work with SDXL today, but the basic idea remains the same.
You need an interface capable of loading the model, a compatible model checkpoint, and hardware that can handle the workload.
For someone coming from Stable Diffusion already, AUTOMATIC1111 is one of the most familiar places to start.
It gives you a traditional web interface with text fields, generation controls, image to image tools, extensions and other features arranged in a relatively straightforward layout.
For someone completely new, that can be considerably easier to understand than a node based interface.
Installing AUTOMATIC1111
The original setup process required a few basic pieces of software.
You generally need Python, Git and suitable GPU support if you plan to generate images locally.
The exact installation requirements can change as the software develops, so it is worth checking the current project documentation before installing anything.
AUTOMATIC1111 Stable Diffusion Web UI on GitHub
The important point is that you do not need to understand the entire technical stack before generating your first image.
Once the interface is installed, you can load a compatible SDXL checkpoint and start experimenting.
For an NVIDIA GPU, CUDA support is part of the wider software stack that allows the generation workload to run efficiently on the GPU.
Your available VRAM also matters considerably.
SDXL is considerably more demanding than older Stable Diffusion checkpoints, especially when you start working with larger resolutions, multiple ControlNets, upscaling and other refinement techniques.
If your system struggles, that does not necessarily mean you have configured everything incorrectly.
Sometimes the hardware simply needs a lighter workflow.
Choosing an SDXL Model
This is one of the areas where beginners can become confused very quickly.
When people say “Stable Diffusion XL,” they can mean the original SDXL base model, a fine tuned checkpoint based on SDXL, or an entire workflow built around one of those models.
The original SDXL release from Stability AI included a Base model and a Refiner model.
The Base model performs the primary generation.
The Refiner can then be used to improve certain details during the later portion of the generation process.
This two model setup was one of the features that initially made SDXL feel more complicated than previous Stable Diffusion workflows.
You suddenly had another model to understand.
Another checkpoint to load.
Another stage in the workflow.
Another thing that could affect generation speed.
For someone learning how to use Stable Diffusion XL, I would not recommend making the Refiner your first concern.
Start with the Base model.
Learn how prompts behave.
Learn how resolution affects composition.
Learn how sampling works.
Learn how image to image behaves.
Then bring the Refiner into the workflow once you understand what you are trying to improve.
That makes the learning process much less confusing.
Your First SDXL Generation
Once SDXL is loaded, resist the temptation to immediately throw a gigantic prompt at it.
Start simple.
Try something like:
A medieval wizard standing inside an ancient stone library, warm candlelight, cinematic fantasy artwork
Generate several images.
Do not judge the model from one generation.
This is one of the most important lessons I learned while experimenting with AI image generation.
A single seed can produce an interesting composition, while another seed with exactly the same prompt can produce something completely different.
That happens because the generation process starts from a random noise pattern.
Your prompt guides the diffusion process, but the initial noise influences the final composition.
So when testing a prompt, generate multiple variations.
You may discover that the prompt itself was perfectly fine and one particular generation simply went sideways.
Understanding Seeds
The seed is one of the most useful controls when you start learning SDXL seriously.
A seed determines the initial noise used for the generation.
When you keep the same prompt and settings but change the seed, you can get different compositions.
When you keep the seed and the other settings the same, you can reproduce a generation much more reliably.
This becomes extremely useful once you find an image that is almost perfect.
Imagine you generate a fantasy character and everything looks excellent except the clothing.
You do not necessarily want to start from scratch.
Keeping the seed and other settings gives you a stable starting point for experimentation.
From there you can change the prompt, use image to image, apply a LoRA, introduce ControlNet or move into an inpainting workflow.
That is where SDXL starts becoming much more than a simple text to image generator.
Understanding Steps, CFG and Sampling
Three settings tend to confuse almost everyone when they first start using Stable Diffusion XL.
Steps.
CFG.
Sampler.
They sound technical, but you can understand the practical effect without becoming a machine learning researcher.
Sampling Steps
The number of steps controls how many iterations the diffusion process goes through while producing the image.
A very low number can produce an image that lacks refinement.
Increasing the steps can give the model more opportunity to develop details, but more steps do not automatically mean better images.
At some point you are simply spending more time for little visible improvement.
For SDXL, a moderate step count is generally a sensible starting point.
You can then test higher and lower values with the same prompt and seed.
That experiment is much more useful than blindly copying somebody else's preferred setting.
CFG Scale
CFG, or classifier free guidance, controls how strongly the generation follows your prompt.
A low value gives the model more freedom.
A high value pushes it harder toward the prompt.
This sounds like higher should be better.
It is not.
Push CFG too far and images can become harsh, unnatural or overly constrained.
A moderate value is usually a better starting point.
The important lesson is that CFG is not a “quality” slider.
It is a guidance control.
That distinction becomes especially important when you start working with complex prompts.
Samplers
The sampler determines how the diffusion process moves from noise toward the final image.
Different samplers can produce noticeably different results even when the prompt, seed, steps and other settings remain unchanged.
This is one reason you may see somebody else's SDXL settings online and fail to reproduce their result.
Their sampler may be different.
Their checkpoint may be different.
Their VAE may be different.
Their LoRA may be different.
Their seed may be different.
Their resolution may be different.
AI image generation has many moving parts.
Do not assume one setting explains the entire result.
Why Your First SDXL Images Can Look Broken
This was one of my biggest frustrations when I started.
You see spectacular SDXL images online and think:
“Why does mine look nothing like this?”
Then you inspect your own output.
The composition looks promising.
The lighting looks good.
The character looks interesting.
Then you notice the hand.
Six fingers.
One finger going through the wrist.
A strange elbow.
An eye slightly detached from the face.
A sword merging into someone's arm.
A background object that makes absolutely no sense.
Welcome to AI image generation.
One of the strangest things about generative imagery is that an image can look excellent from several feet away and completely ridiculous when you zoom in.
I started thinking of this as the Where's Waldo problem of AI art.
You initially see a beautiful image.
Then you start searching for everything that went wrong.
The problem is that a diffusion model does not have the same understanding of anatomy, physics and object permanence that a human artist does.
It has learned visual patterns.
Sometimes those patterns produce something remarkably convincing.
Sometimes they produce something that looks convincing until you inspect the details.
This is why refinement tools become so important.
LoRAs: One of the Most Useful SDXL Tools
Once you become comfortable with basic generation, LoRAs are probably one of the first advanced concepts worth learning.
LoRA stands for Low Rank Adaptation.
You can think of a LoRA as a relatively small add on that modifies the behavior of a larger model.
Rather than replacing the entire checkpoint, it adds learned information that can push the generation toward a particular concept, character, style, clothing type, visual aesthetic or other subject.
For example, imagine you want a particular fantasy illustration style.
You could search for an SDXL compatible LoRA trained around that style.
Load your primary checkpoint.
Load the LoRA.
Add its trigger words to your prompt when required.
Then adjust its strength.
That gives you another layer of control over the image.
LoRAs become particularly interesting when you are building recurring characters.
Suppose your Dungeons & Dragons campaign has a character who needs to appear in multiple scenes.
A carefully trained character LoRA can help the model reproduce recognizable characteristics across different generations.
It will not guarantee perfect consistency, but it can make the problem considerably easier.
That becomes even more valuable when you eventually move toward animation and video workflows.
ADetailer and Fixing Faces and Hands
One of the first things you notice with AI generated people is that faces and hands can require additional work.
This is where tools such as ADetailer became popular in Stable Diffusion workflows.
ADetailer detects areas such as faces or hands and runs additional refinement on those regions.
The idea is simple.
Generate the complete image first.
Detect the problematic region.
Process that region separately.
Place the refined result back into the original image.
This can be remarkably useful.
Imagine generating a fantasy warrior.
The armor looks fantastic.
The environment is perfect.
The lighting works.
The pose is exactly what you wanted.
But the character's face has strange eyes and one hand has an impossible number of fingers.
Throwing the entire image away would be frustrating.
A detailer workflow gives you a chance to repair those areas without regenerating everything.
There is a catch.
Automatic refinement can also introduce changes you did not want.
A face may become too generic.
Skin texture may change.
Hair can be altered.
The hand can improve while the surrounding background gets slightly damaged.
So treat these tools as another generation stage rather than a guaranteed repair button.
ControlNet: When You Need More Control
This is where Stable Diffusion XL becomes particularly interesting.
Text prompts are excellent for describing what you want.
They are much less reliable at describing exactly where everything should go.
Suppose you have a photograph of someone standing in a specific pose.
You want to transform that person into a fantasy character.
A prompt can describe the fantasy character.
It cannot easily tell the model precisely how the arms, legs, torso and head should be positioned.
ControlNet can help.
ControlNet allows an additional image based signal to guide the generation.
Instead of saying:
A warrior standing with one hand raised
you can provide an image containing the pose you want.
The model then uses information extracted from that image as an additional structural guide.
This is incredibly useful for character creation.
It is also one of the reasons ControlNet becomes important once you move beyond simple experimentation and start creating assets for a project.
OpenPose ControlNet
OpenPose is designed around human body positioning.
It identifies key points within a person's pose and creates a representation of that structure.
You can then feed that information into your generation workflow.
Imagine you have an image of someone waving.
You want to transform that person into an elf.
The original image gives you the pose.
Your prompt describes the elf.
The ControlNet provides structural guidance.
The result can preserve much of the original pose while changing the character, clothing, environment and overall visual style.
For a D&D project, this can be particularly useful.
You could create reference poses for your characters and reuse them across scenes.
The pose can remain relatively stable while the visual content changes around it.
Depth ControlNet
Depth provides another kind of structural information.
Rather than concentrating on skeletal pose, it estimates the depth relationships within an image.
Foreground objects appear closer.
Background objects appear farther away.
The resulting depth map gives the generation a rough understanding of spatial arrangement.
This can help when you care about the composition of an existing image.
If the original image contains a character standing in front of a castle, the depth information can help preserve the relationship between foreground and background while you change the actual visual content.
The strength matters considerably.
Push the influence too high and the generation can become overly constrained.
Keep it lower and SDXL has more freedom to reinterpret the scene.
Finding that balance is largely a matter of experimentation.
Canny ControlNet
Canny takes a different route.
It extracts edges from an image.
Those edges can then guide the generated image.
This is useful when you care about the basic structure of an image but want the model to reinterpret its visual content.
Imagine taking a photograph of a person standing beside a car.
Canny can capture the outlines.
You can then prompt SDXL for a cyberpunk character standing beside a futuristic vehicle.
The structural information remains useful while the model changes the visual interpretation.
The important thing to remember is that ControlNet strength determines how much freedom the model has.
A high value can preserve too much of the original structure.
A low value gives SDXL considerably more room to reinterpret the scene.
There is no universal perfect value.
The right setting depends on how much of the original image you want to preserve.
Inpainting: My Favorite Way to Fix an Image
If there is one feature I repeatedly return to when working with generated images, it is inpainting.
The concept is beautifully simple.
You have an image you mostly like.
Something is wrong.
You mask the problematic area.
You provide a new prompt.
The model generates a replacement for that region while attempting to blend it into the surrounding image.
Imagine you have generated a wizard standing on a mountain.
Everything looks great except the wizard is holding a sword that looks like a melted piece of metal.
You do not need to regenerate the mountain, sky, lighting and entire character.
Mask the sword.
Describe the replacement.
Generate several versions.
Pick the one that fits.
This is much closer to how I like to think about AI image generation.
The first generation is not necessarily the final image.
It can be the foundation.
You gradually correct the parts that need attention.
Outpainting
Outpainting works in the opposite direction.
Instead of replacing an existing area, you expand the canvas and generate additional content outside the original boundaries.
This can be useful when an image has a great subject but poor framing.
Imagine generating a character portrait that looks fantastic but is too tightly cropped.
You can extend the canvas.
Give the model information about the surrounding environment.
Then generate the missing area.
This can be useful for wallpapers, cinematic compositions, posters and assets that need to fit different aspect ratios.
It is also valuable when you create artwork that later needs to be adapted for different platforms.
A portrait created for one format can sometimes be expanded into a wider composition without recreating the entire scene.
Prompting Stable Diffusion XL
Prompting is probably the part of AI image generation that receives the most attention.
There is a reason for that.
Your prompt matters.
But prompt writing is often misunderstood.
A longer prompt does not automatically produce a better image.
You want to communicate the important visual information clearly.
For example:
A young elven ranger standing in an ancient forest, leather armor, bow across her back, morning fog, warm sunlight passing through tall trees, cinematic fantasy illustration, highly detailed environment
This gives the model several useful concepts:
- Character
- Location
- Clothing
- Object
- Lighting
- Atmosphere
- Style
- Detail level
That is usually more useful than stuffing dozens of unrelated quality terms into the prompt.
Prompt Weighting
One of the useful features associated with Stable Diffusion interfaces is prompt weighting.
You can increase the influence of a particular concept.
For example:
(red cloak:1.3)
tells the interface to give that concept greater emphasis.
You can reduce influence with a value below 1.
For example:
(red cloak:0.7)
The exact behavior depends on the interface and parser being used, so it is worth checking the syntax supported by your current setup.
The practical lesson is simple.
If an important concept keeps disappearing, you can give it more emphasis.
If something keeps dominating the image, reduce its influence.
Do not push weights too aggressively.
Very strong weighting can produce strange compositions because you are effectively asking the model to prioritize one concept over everything around it.
Prompt Order Can Matter
The ordering of concepts can also influence results.
Consider these two prompts:
A man wearing a red coat sitting inside a quiet cafe
and
A man sitting inside a quiet cafe wearing a red coat
They describe almost the same thing.
The model can still interpret them differently.
This becomes more noticeable with complicated prompts containing several subjects, visual styles and environmental details.
I generally prefer putting the most important information toward the beginning of the prompt.
Start with the subject.
Then describe the action or composition.
Then add clothing and important objects.
Then describe the environment.
Then lighting and visual style.
This gives the prompt a natural structure and makes it easier to modify later.
Prompt Editing and Concept Mixing
Once you become comfortable with basic prompting, you can experiment with more advanced prompt syntax.
Prompt editing can change one concept into another during the generation process.
For example, a workflow can begin with one concept and later transition toward another.
This can produce unusual combinations because the image structure begins forming around one idea before the model receives a different concept.
It can be especially fun for hybrid creatures and fantasy concepts.
You might experiment with something like:
A [wolf:dragon:0.5]
The idea is that the generation begins with one concept and transitions toward another during the process.
The exact syntax and behavior depend on the interface and parser.
This is one area where experimentation is far more useful than memorizing a collection of tricks.
Why ComfyUI Keeps Coming Up
If you have spent any time looking into SDXL, you have probably encountered ComfyUI.
And if you opened it for the first time, you may have immediately thought:
“What on earth am I looking at?”
I had the same reaction.
ComfyUI is a node based interface.
Rather than presenting every operation as a button or tab, it represents the image generation process as connected nodes.
That can look intimidating.
Once you understand the concept, however, it becomes easier to appreciate why people use it.
A traditional interface hides much of the workflow.
ComfyUI exposes it.
- You can see the model loading.
- You can see the positive prompt.
- You can see the negative prompt.
- You can see the sampler.
- You can see the latent image.
- You can see the VAE.
- You can see ControlNet.
- You can see the image being decoded.
Then you can connect additional components wherever they make sense.
For complex SDXL workflows, this can be incredibly powerful.
But you do not need to learn every node.
That is probably the biggest mistake beginners make.
Start with a working workflow.
Change one thing.
Generate.
Observe the result.
Then add another component.
Why ComfyUI Is Useful for SDXL
One of the things I particularly like about node based workflows is that you can keep multiple stages together.
Instead of manually switching between different tools, you can create a workflow that handles several operations in sequence.
For example:
SDXL checkpoint → prompt → sampler → ControlNet → detail refinement → upscale → final image
Once that workflow is configured, you can reuse it.
You can also save the workflow along with generated images in supported formats.
That means you can return to an image later and inspect how it was generated.
This is one of the most useful aspects of ComfyUI for learning.
Someone can share an image containing the workflow metadata.
You can load that image.
The workflow appears.
Suddenly a complicated generation becomes something you can inspect and modify.
You do not have to recreate every connection manually.
Where SDXL Starts Becoming Interesting for D&D
This is the point where my interest in Stable Diffusion XL becomes much more practical.
Generating random fantasy artwork is fun.
Creating assets for an actual D&D campaign is a different problem.
- You need consistency.
- You need recurring characters.
- You need locations that feel like they belong to the same world.
- You may need portraits, monsters, environments, battle scenes, maps and cinematic moments.
A simple text to image workflow can produce individual images.
A more advanced SDXL workflow can become a small visual production system.
For example, you could create a character reference image.
Then use a character LoRA to reproduce that character.
- Use OpenPose when you need specific body positioning.
- Use ControlNet when you need to preserve composition.
- Use inpainting when the hands or face need correction.
- Use outpainting when the image needs a wider composition.
Then use an upscaling stage for the final asset.
Suddenly you are no longer asking SDXL to magically produce a finished image.
You are directing the image synthesis process.
That difference makes the system much easier to work with.
From Stable Diffusion XL to AI Video
If your long term goal is AI video for your D&D campaign, learning SDXL is still useful.
Video generation introduces another problem that does not exist to the same degree with individual images.
Consistency across time.
A single image can look perfect.
A sequence of images can look terrible if the character's face, clothing, proportions, lighting or environment changes between frames.
This is why I would spend time learning image generation before jumping headfirst into video.
- Learn how to create a consistent character.
- Learn how to control poses.
- Learn how to preserve environments.
- Learn how to repair individual frames.
- Learn how seeds affect results.
- Learn how reference images influence generation.
- Learn how LoRAs and ControlNet affect consistency.
Those concepts become valuable when you start working with moving images.
The tools may change, but the underlying creative problem remains similar.
You are trying to tell a generative model what should remain stable and what is allowed to change.




