AI Image Settings Explained: Seed, Steps, CFG and Samplers
Learn what seed, inference steps, CFG guidance, samplers and schedulers do in AI image generation, why settings change results, and how to tune them.
Approximately 21 min read
Open an AI image generator and you may see settings such as:
Seed
Steps
CFG Scale
Sampler
Scheduler
Width
Height
Negative Prompt
Denoising Strength
At first, these controls can look like mysterious numbers that need to be copied from someone else’s workflow.
They are not.
Each setting controls a different part of the generation process.
A useful mental model is:
Model
+
Prompt
+
Initial randomness
+
Sampling process
+
Guidance
+
Resolution
=
Generated image
The most important settings to understand first are:
Seed
Steps
CFG
Sampler / Scheduler
Once you know what those four controls actually do, image-generation interfaces become much easier to use.
Start with the basic image-generation process
Many diffusion-style image generators begin from random noise.
Conceptually:
Random noise
↓
Denoising step
↓
Denoising step
↓
Denoising step
↓
...
↓
Final image
The text prompt conditions this process so that the denoising moves toward visual content related to your request.
For example:
a red fox sitting beside a lake at sunrise
provides text conditioning.
But the final result also depends on:
which random noise you started with
how many updates were performed
which update algorithm was used
how strongly the prompt influenced generation
Those correspond roughly to:
seed
steps
sampler / scheduler
CFG
For the underlying architecture, read How Does AI Image Generation Work? Diffusion Models Explained.
What is a seed?
AI image generation normally includes randomness.
The initial latent noise is generated using a pseudo-random number generator.
The:
seed
controls the initial random state.
Conceptually:
Seed 12345
↓
Pseudo-random number generator
↓
Noise pattern A
↓
Image generation
Change the seed:
Seed 98765
and you get:
Noise pattern B
That different starting point can produce a dramatically different composition even if every other setting remains unchanged.
Same prompt, different seed
Suppose your prompt is:
a cabin beside a frozen lake,
mountains in the distance,
cinematic winter photography
Generate with:
Seed 100
and you might get:
cabin on left
lake in center
mountains behind
Generate with:
Seed 101
and you might get:
cabin centered
larger mountains
different camera angle
different clouds
The prompt stayed the same.
The starting noise changed.
A seed is not an image ID
A common misunderstanding is:
Seed 1234
=
one specific image stored somewhere
That is not correct.
The seed initializes a random-number generator.
The actual output also depends on:
model
model version
prompt
negative prompt
scheduler
steps
CFG
resolution
software
precision
other pipeline settings
So the seed is only one part of the generation configuration.
Same seed does not always mean identical pixels
If you reproduce:
same model
same prompt
same seed
same scheduler
same steps
same CFG
same resolution
you can often obtain very similar or reproducible results within the same software environment.
But exact cross-system reproducibility is harder.
Differences can come from:
GPU hardware
CPU vs GPU
PyTorch version
library version
precision
kernel implementation
deterministic settings
So:
same seed
≠
universal image identifier
For serious reproducibility, record the complete environment.
Why seeds are useful
Seeds are valuable for three main reasons.
Exploration
Generate many seeds:
100
101
102
103
104
until you find a composition you like.
Refinement
Once you find a good seed, keep it fixed while adjusting:
prompt wording
CFG
steps
sampler
This reduces one source of variation.
Testing
When comparing two settings, using the same seed makes the experiment more controlled.
For example:
Seed: 12345
Steps: 20
vs
Seed: 12345
Steps: 30
Now the initial randomness is similar, so the step count is easier to compare.
Random seed mode
Many interfaces offer:
Seed: -1
or:
Random Seed
This generally means:
choose a new seed automatically
for each generation.
That is useful during exploration.
Once you find a result worth refining, save the actual seed used.
What are inference steps?
Diffusion-style generation is iterative.
The model does not normally transform pure noise directly into the finished image in one ordinary denoising call.
Instead, the sampling process follows a sequence of noise levels or timesteps.
For example:
Start
↓
Step 1
↓
Step 2
↓
Step 3
↓
...
↓
Step 20
↓
Final image
The interface may call this:
Steps
Sampling Steps
Inference Steps
Denoising Steps
What happens during one step?
At a simplified level:
Current noisy latent
↓
Denoising model predicts useful information
↓
Scheduler / sampler applies an update
↓
New latent
Then the process repeats.
The model prediction and scheduler work together.
The model does not independently decide the entire mathematical trajectory.
More steps usually mean more computation
Suppose one generation uses:
20 steps
and another:
40 steps
The 40-step version normally requires substantially more denoising work.
So, all else equal:
more steps
→ longer generation time
But:
more steps
≠
automatically better image
This is one of the most important rules in image generation.
Why more steps stop helping
The purpose of sampling is to move the noisy representation toward a useful final sample.
After enough effective updates, additional steps may provide diminishing returns.
Conceptually:
1 step
→ insufficient
5 steps
→ rough
20 steps
→ strong result
40 steps
→ perhaps modest difference
100 steps
→ possibly little useful improvement
The exact behavior depends on:
model
scheduler
distillation
noise schedule
guidance
prompt
Some models are explicitly designed to work with very few steps.
Others expect more.
There is no universal number.
Why “30 steps is best” is bad advice
You may see recommendations such as:
Always use 30 steps.
That recommendation lacks context.
For one pipeline:
4 steps
may be appropriate.
Another may work well around:
20–30
Another scheduler may need a different range.
The correct question is:
How many steps does this specific
model + scheduler need?
Check the model documentation first.
Distilled models can use fewer steps
Some image models are trained or adapted specifically for fast generation.
Their intended workflow may use dramatically fewer steps than older diffusion models.
If a model was designed for:
4–8 steps
forcing:
50 steps
does not automatically improve it.
The architecture and training objective matter.
This is why copying Stable Diffusion 1.5 settings into every modern image model is increasingly unreliable.
What is a sampler?
Different interfaces use the words:
sampler
and:
scheduler
somewhat differently.
At a high level, the sampler/scheduler determines how the current noisy representation is updated during generation.
Conceptually:
Model prediction
+
Current latent
+
Current noise level
↓
Sampling algorithm
↓
Next latent
Different algorithms follow different numerical paths toward the final image.
Scheduler vs sampler
In Hugging Face Diffusers, the term:
scheduler
is widely used.
Other image-generation interfaces frequently say:
sampler
The terminology is not perfectly standardized.
For practical purposes, both terms often refer to machinery controlling the iterative denoising trajectory.
However, specific software may distinguish:
solver algorithm
noise schedule
sigma schedule
sampler
scheduler
more precisely.
So always interpret the name within the software you are using.
The scheduler is not the trained model
Suppose you use the same checkpoint:
Model X
and switch from one scheduler to another.
You did not retrain or replace the model.
You changed the numerical sampling procedure used during inference.
Conceptually:
same trained model
+
different route through denoising
=
potentially different output
That is why scheduler choice can change image appearance even when the seed and prompt remain fixed.
Why different samplers look different
Imagine several possible paths between:
random noise
and:
finished sample
One solver might take:
many conservative updates
while another uses:
fewer more efficient updates
Some include stochastic behavior.
Others can be more deterministic.
The paths are mathematically different, so final images can differ.
DDPM
Denoising Diffusion Probabilistic Models established the classic iterative diffusion framework.
The reverse process involves many denoising transitions.
High sample quality was possible, but traditional sampling could require many model evaluations.
That motivated work on faster sampling algorithms.
DDIM
DDIM stands for:
Denoising Diffusion Implicit Models
DDIM showed that models trained using the diffusion objective could be sampled through a different, non-Markovian process that can require fewer steps.
This is historically important because it demonstrated that:
training process
and:
inference sampling trajectory
do not have to be identical.
Modern diffusion software now exposes many solver and scheduler choices beyond DDIM.
Euler
You may see:
Euler
or:
Euler Ancestral
Euler a
in image-generation interfaces.
These are numerical sampling approaches inspired by differential-equation solvers.
They follow different update rules.
Euler Ancestral includes stochastic characteristics that can produce variation even within related configurations.
Do not assume an Euler result will exactly match a DDIM or DPM-family result at the same seed.
DPM and DPM++
You may also encounter:
DPM
DPM++
DPM++ 2M
DPM++ SDE
These represent solver families designed for efficient diffusion sampling.
Different variants can balance:
step efficiency
stochasticity
detail
stability
differently.
The exact best choice depends on the model.
Flow-matching models use related but different schedulers
Modern image generation increasingly includes:
flow matching
rectified flow
architectures.
These are not simply classic DDPM sampling under a different label.
Libraries therefore expose schedulers such as:
FlowMatchEulerDiscreteScheduler
FlowMatchHeunDiscreteScheduler
for appropriate flow-based pipelines.
This is another reason not to assume that one Stable Diffusion sampler recommendation applies universally.
What is a noise schedule?
The generation process traverses different noise levels.
The:
noise schedule
defines those levels.
Conceptually:
very noisy
↓
less noisy
↓
less noisy
↓
nearly clean
But the steps do not necessarily need to be spaced evenly.
A scheduler may concentrate more updates in portions of the trajectory where they are most useful.
What are sigmas?
Some diffusion and flow implementations express noise levels using:
sigma
σ
A sequence may look conceptually like:
large sigma
↓
medium sigma
↓
small sigma
↓
0
Different sigma schedules can change how sampling effort is distributed.
This means two workflows with:
20 steps
may not evaluate the same 20 noise levels.
The scheduler configuration matters.
Karras sigmas
Some schedulers support:
Karras sigmas
which use a particular arrangement of noise levels.
These became popular in diffusion interfaces because they can provide efficient sampling for compatible models and solvers.
But they should not be treated as a universal:
Karras = better
switch.
Use them when appropriate for the model and scheduler.
What is CFG?
CFG stands for:
Classifier-Free Guidance
It controls how strongly generation is guided toward the conditioning prompt in compatible diffusion pipelines.
A simplified intuition is:
low CFG
→ more freedom from prompt
higher CFG
→ stronger prompt pressure
But that does not mean:
higher CFG
=
better
Why classifier-free guidance exists
Conditional diffusion wants the generated sample to correspond to information such as:
text prompt
Classifier-Free Guidance combines conditional and unconditional model predictions to strengthen the effect of conditioning.
Conceptually:
Prediction without prompt pressure
+
Prediction conditioned on prompt
↓
Guidance
↓
Stronger conditional direction
The original CFG work describes this as a trade-off involving sample fidelity and diversity.
What is CFG scale?
Interfaces may display:
CFG Scale: 7
or:
Guidance: 5
Higher values increase the influence of text conditioning in pipelines that use CFG.
A simplified view is:
CFG 1
→ weak/no extra CFG effect
CFG 5
→ stronger conditioning
CFG 10
→ stronger again
But these numbers have no universal meaning across all model families.
Why very high CFG can look bad
Excessive guidance can push the sample too aggressively toward prompt-conditioned predictions.
Depending on the model, this can contribute to effects such as:
oversaturation
harsh contrast
unnatural edges
artifacts
reduced realism
So guidance trades:
prompt adherence
against other aspects of image quality and diversity.
Higher is not automatically better.
Different models expect different CFG ranges
A Stable Diffusion-era workflow might use one range.
A newer transformer or flow-based model may expect:
different guidance
or use CFG differently.
Some pipelines can even be designed to work with little or no ordinary classifier-free guidance.
Therefore:
CFG 7 is the best setting
is not a reliable universal rule.
Use the model’s recommended configuration.
CFG and prompt strength are different ideas
If an image ignores your prompt, increasing CFG is not always the correct fix.
The problem could instead involve:
model capability
poor prompt structure
unsupported concept
text encoder limitations
too many conflicting instructions
wrong checkpoint
Guidance cannot force a model to represent knowledge it does not have.
What is a negative prompt?
Many CFG-based image pipelines support:
negative prompt
A positive prompt says:
move generation toward these concepts
while a negative prompt provides conditioning representing concepts to move away from.
Example:
Prompt:
studio portrait, natural skin, soft lighting
Negative:
blurry, low resolution
The negative prompt participates in the guidance process.
Giant negative prompts are not automatically better
You may find copied negative prompts containing dozens or hundreds of words:
bad anatomy,
bad hands,
ugly,
blurry,
deformed,
...
These can sometimes help a particular model or workflow.
But they are not universal quality spells.
Long negative prompts can change conditioning in unintended ways.
Newer model families may respond very differently from older Stable Diffusion checkpoints.
Use negative prompts deliberately.
Negative prompts may be irrelevant without CFG
In some diffusion pipelines, negative prompts are used through the guidance mechanism.
If ordinary classifier-free guidance is disabled, the negative prompt may not have the same effect or may be ignored by that pipeline.
Again, interface behavior depends on the implementation.
Seed and CFG interact
Suppose:
Seed: 100
CFG: 4
produces a composition you like.
Changing only:
CFG: 9
can alter more than colors.
Because each sampling update changes, later latent states diverge.
The final image may differ in:
composition
details
lighting
object shape
even though the initial noise was identical.
A fixed seed stabilizes the starting point, not the entire image.
Seed and sampler interact
Likewise:
Seed 100
+
Euler
and:
Seed 100
+
DDIM
should not be expected to produce identical images.
The same initial noise passes through different numerical update rules.
The output can diverge rapidly.
Seed and resolution interact
Suppose you generate:
512 × 512
then change to:
1024 × 1024
while keeping the same seed.
The initial noise tensor has a different shape.
You should not expect the larger image to simply be:
the same composition with more pixels
Resolution is part of the generation configuration.
What does width and height affect?
Resolution controls the spatial size of the generated image.
Pixel count is:
width × height
For example:
512 × 512
=
262,144 pixels
1024 × 1024
=
1,048,576 pixels
That is four times as many output pixels.
Latent diffusion reduces the computational burden by working in compressed representations, but higher resolution still increases memory and computation.
Native model resolution matters
Models are trained using particular image-size distributions.
Generating far outside the model’s normal operating range can produce:
duplicated subjects
strange composition
repetition
poor anatomy
unexpected framing
Modern models may support flexible resolutions and aspect ratios, but you should still follow model-specific recommendations.
What is denoising strength?
Image-to-image workflows introduce another important setting:
denoising strength
or simply:
strength
Instead of starting from pure random noise, img2img starts from an encoded source image and adds noise before denoising.
Conceptually:
Source image
↓
Encode
↓
Latent
↓
Add noise according to strength
↓
Denoise
↓
Modified image
Low denoising strength
A lower value generally means:
less noise added
↓
generation remains closer
to source image
This can help when you want to preserve:
composition
pose
identity
layout
while changing details.
High denoising strength
Higher strength allows the model to move farther from the source.
Conceptually:
more noise
↓
less source information preserved
↓
more freedom to regenerate
At sufficiently high strength, the result may differ dramatically from the original image.
Denoising strength is not CFG
These two settings control different things.
Strength
→ how far img2img moves away from source image
CFG
→ how strongly text conditioning guides generation
Changing one does not substitute for the other.
How should you tune image settings?
Do not randomly change everything.
Use controlled experiments.
Suppose you have:
Prompt:
a small cabin beside a lake at sunrise
Seed:
12345
Steps:
20
CFG:
5
Scheduler:
A
Start there.
Then change one variable.
Test steps
Keep:
seed
prompt
CFG
scheduler
resolution
constant.
Generate:
10 steps
20 steps
30 steps
Compare.
You can now see whether additional inference steps matter for that model.
Test CFG
Return to one step count.
Keep everything else constant.
Try:
CFG 3
CFG 5
CFG 7
if those values are sensible for your model.
Compare:
prompt adherence
color
contrast
composition
artifacts
Do not assume the highest value wins.
Test schedulers
Now keep:
seed
prompt
steps
CFG
constant and switch the scheduler.
For example, where compatible:
Euler
DDIM
DPM-family scheduler
Observe whether the model changes:
detail
composition
texture
stability
Remember that equal step counts do not necessarily imply equal computational or numerical behavior between schedulers.
Then explore seeds
Once your general configuration is good:
Seed 1
Seed 2
Seed 3
Seed 4
...
becomes an efficient way to explore different compositions.
When one looks promising, lock the seed and refine the prompt.
A practical workflow
A strong workflow is:
1. Use model-recommended defaults
2. Write a clear prompt
3. Generate several random seeds
4. Choose the best composition
5. Lock the seed
6. Adjust prompt details
7. Test CFG only if prompt adherence needs tuning
8. Test steps if quality appears underdeveloped
9. Compare schedulers only when there is a reason
10. Save the full generation metadata
This is far more efficient than constantly changing every control.
Save your generation metadata
If you want to recreate an image, record:
model
model revision
prompt
negative prompt
seed
steps
CFG
scheduler
resolution
LoRAs
ControlNet settings
VAE
software version
For highly precise reproducibility, also record:
framework version
hardware
precision
A seed by itself is not enough.
Why model version matters
Suppose a publisher updates a checkpoint.
You use:
same prompt
same seed
same settings
but the model weights changed.
The output may be completely different.
Therefore serious workflows should identify the exact checkpoint or model revision used.
Why LoRA changes the result
A LoRA modifies model behavior.
So:
base model
+
LoRA A
is not the same generation system as:
base model only
or:
base model
+
LoRA B
If reproducing an image, record:
LoRA name
version
weight/strength
as well.
Why VAE can matter
In latent-diffusion architectures, the VAE decodes latent representations into pixels.
Changing the VAE can alter:
color
contrast
fine detail
artifacts
So it can affect final appearance even when the seed and sampling configuration remain unchanged.
Why precision can matter
Inference may run using numerical types such as:
FP32
FP16
BF16
Different hardware and kernels can introduce small numerical differences.
Because diffusion is iterative, small differences can propagate through many steps.
Therefore:
same seed
does not guarantee absolute numerical identity across every system.
Should you chase exact reproducibility?
It depends on the task.
For creative exploration:
No.
Small differences may not matter.
For:
benchmarking
regression tests
workflow debugging
scientific comparisons
reproducibility matters much more.
Then you should control:
software
hardware
seed
model revision
settings
as tightly as possible.
Common mistake: using too many steps
If a model produces strong images at:
20 steps
generating at:
100 steps
may simply waste time.
Test whether the extra computation produces a meaningful benefit.
Do not equate:
more computation
with:
better output
Common mistake: maximizing CFG
Users sometimes see:
CFG 7
and think:
CFG 20 must follow the prompt even better.
Extreme guidance can push generation into poor visual regions.
The correct value is model-dependent.
Start near the model’s recommended setting.
Common mistake: changing seed while tuning
Suppose you compare:
Image A
Seed 100
CFG 5
Image B
Seed 847291
CFG 8
and prefer B.
Was that because:
CFG changed?
or:
seed changed?
You cannot know.
When tuning a parameter, fix the seed.
Common mistake: comparing samplers at different step counts
Suppose:
Sampler A:
15 steps
Sampler B:
40 steps
Then conclude:
Sampler B has better detail.
That may simply reflect the different computational budget.
For cleaner comparison, control as many settings as practical.
Common mistake: importing old Stable Diffusion settings into every new model
Image-model architectures are evolving rapidly.
Advice that worked for:
Stable Diffusion 1.5
does not automatically transfer to:
SDXL
SD3
flow-matching models
distilled models
other image transformers
Model documentation should override generic internet recipes.
Common mistake: calling every scheduler a quality upgrade
Different schedulers are numerical methods.
There is no universal ranking such as:
Sampler X
>
Sampler Y
>
Sampler Z
for every model, prompt, step count, and aesthetic.
The appropriate scheduler is pipeline-dependent.
A simple beginner baseline
If you are using an unfamiliar image model:
1. Use the model's default scheduler
2. Use the recommended step count
3. Use the recommended guidance setting
4. Generate random seeds
5. Find a composition you like
6. Lock the seed
7. Tune one setting at a time
This avoids most unnecessary complexity.
Advanced users should still benchmark settings systematically
More experience does not make controlled testing less important.
For a specific model, create a grid using:
same prompts
same seeds
different step counts
different CFG values
different schedulers
Then compare:
image quality
prompt adherence
generation time
VRAM
failure patterns
This gives you model-specific evidence rather than folklore.
Sampler speed vs image quality
A scheduler that produces acceptable results at:
10 steps
may be more useful than one requiring:
40 steps
even if the 40-step image is marginally better.
For real workflows, optimize:
quality per unit time
not only maximum theoretical quality.
Batch generation
Suppose you want four concepts.
Instead of heavily tuning one seed immediately:
generate several seeds
↓
compare compositions
↓
select candidate
↓
refine candidate
This often produces better creative results than endlessly manipulating CFG or step count on an unpromising composition.
Seeds are excellent for controlled A/B tests
Suppose you want to know whether a new LoRA helps.
Run:
same model
same prompt
same seed
same scheduler
same steps
same CFG
with:
LoRA off
and:
LoRA on
Now the difference is easier to interpret.
The same principle applies to:
VAE
ControlNet
scheduler
CFG
prompt edits
What if an image looks almost right?
Do not automatically change the seed.
If composition is good but details are wrong, try small changes first:
prompt refinement
negative prompt refinement
LoRA strength
ControlNet
inpainting
Changing the seed may destroy the composition you already like.
What if composition is completely wrong?
Then changing the seed is often appropriate.
The seed strongly influences the initial latent structure, so exploring several seeds is a fast way to search different layouts.
Think:
wrong composition
→ explore seeds
good composition, wrong details
→ keep seed and refine
This is a very useful practical rule.
What if the prompt is ignored?
Check:
Does the model understand the concept?
Is the prompt contradictory?
Is CFG appropriate?
Is the model intended for this type of prompt?
Is a LoRA overriding behavior?
Is the text encoder overloaded by a very long prompt?
Do not automatically increase CFG to the maximum.
What if the image looks oversaturated or harsh?
Possible causes include:
excessive guidance
model-specific aesthetic
VAE
prompt language
post-processing
If CFG is unusually high, reduce it and compare using the same seed.
What if images look unfinished?
Possible causes include:
too few steps
wrong scheduler for model
incorrect model configuration
very aggressive fast-generation settings
Increase steps gradually and compare.
If improvement stops, additional steps are unlikely to solve the problem.
What if changing scheduler completely changes composition?
That is normal.
The scheduler controls the trajectory through latent space.
Same:
seed
+
prompt
does not force different schedulers to follow the same path.
Treat sampler choice as part of the complete generation recipe.
Four settings to remember
If you are new to image generation, remember these definitions:
Seed
→ starting randomness
Steps
→ number of sampling updates
CFG
→ strength of prompt guidance in compatible pipelines
Sampler / Scheduler
→ numerical procedure used to update the sample
Do not collapse them into one generic:
quality setting
They control different mechanisms.
The settings work together
Think of generation as:
Seed
↓
creates starting noise
Prompt + CFG
↓
define conditioning pressure
Model
↓
predicts how the noisy sample should change
Scheduler
↓
applies numerical update
Steps
↓
determine how many updates occur
VAE / decoder
↓
produces final pixels
Changing any part can change the final image.
Bottom line
AI image settings are easier to understand when you stop treating them as magic numbers.
Seed controls the initial random state:
same seed
→ similar starting noise
but exact reproducibility also depends on the rest of the pipeline.
Steps control how many inference updates are performed:
more steps
→ more computation
but not necessarily better quality.
CFG controls prompt guidance in compatible classifier-free-guidance pipelines:
higher CFG
→ stronger prompt conditioning
but excessive guidance can reduce visual quality.
Sampler or scheduler controls the numerical path through the denoising process:
same model
+
different scheduler
→ potentially different image
The practical workflow is:
start with model defaults
↓
explore seeds
↓
lock a promising seed
↓
change one parameter at a time
↓
measure whether it actually improves the image
There is no universal:
best seed
best CFG
best sampler
best number of steps
for every AI image model.
The correct settings depend on the model, scheduler, workload, and desired result.
Once you understand what each control does, you can stop copying random parameter recipes and start tuning image generation systematically.