DIGITAL DIVAS

AI Image Generation

Creating Stunning AI Influencers: The Complete Stable Diffusion Guide for 1.5, SDXL, and FLUX

By Digital Divas Team · March 22, 2025 · 44 min read

3D render of a cute robot character holding a paintbrush and standing next to a canvas
3D render of a cute robot character holding a paintbrush and standing next to a canvas

Great AI influencers start with precise prompts. Whether you're creating a photorealistic influencer, anime-inspired personalities, or cinematic-style images, clear and detailed prompts are crucial.

Stable Diffusion 1.5: Offers vast creative flexibility, excelling particularly with artistic styles like anime or conceptual visuals. For a deep dive on getting the most out of this model specifically, see our deep dive on Stable Diffusion 1.5 optimization.

Stable Diffusion XL (SDXL): Best for ultra-high-quality, realistic influencer images thanks to its enhanced detail and resolution.

FLUX.1: Ideal for influencers requiring sophisticated details, readable text, and accurate anatomy through natural language prompts.

1. Understanding the Models and Their Differences

Before crafting prompts, it's important to know how SD 1.5, SDXL, and FLUX differ under the hood. These differences affect how you should prompt them:

Key Differences Summary:

ModelText Encoder(s)Prompt StyleSupports Prompt Weights?StrengthsWeaknesses / Notes
SD 1.5CLIP (ViT-L/14, 75-token limit)Often tag-based or short phrases (e.g. masterpiece, best quality, 1girl, ...). Can also use descriptive sentences.Yes (with (), [], or :n syntax in A1111)Huge variety of custom models (anime, photorealistic, etc.). Good with prompt weighting and negative prompts to fine-tune results.Struggles with text (signs, letters). Needs more negative prompts to avoid artifacts. Native 512px can show deformities at higher res.
SDXL 1.02× CLIP (OpenCLIP ViT-G + CLIP ViT-L)Best with a descriptive sentence + style tags (e.g. "A portrait of a warrior princess in a forest, intricate detail, 8k, photorealistic"). Handles longer prompts well.Yes (same syntax as SD1.5). Also allows using two prompts (one per encoder) in some UIs.Higher native resolution (1024px), better detail & prompt recognition. Dual encoders can improve fidelity. Refiner model can enhance details after initial image.Still not great at text (signs/logos). Requires more VRAM. Benefits from negative prompts for best results. Limited fine-tuned models compared to 1.5 (as it's newer).
FLUX.1CLIP-L + T5-XXL (dual encoding)Best with natural language prompts (as if explaining to a human). e.g. "A futuristic city skyline at sunset, with neon signs reflecting on wet streets." Avoid complex weight syntax (not needed).Not in the same way — FLUX doesn't use () weights; use phrasing like "with emphasis on…" instead.Excellent prompt adherence; often no need for heavy prompt "hacks." Renders legible text and complex compositions well. Great anatomy (hands, faces) out of the box. Good results with simple prompts.

Tip: If you're deciding which model to use for a task, consider the subject:

Next, we'll cover how to craft prompts (positive and negative) effectively for each model, and then we'll get into using the interfaces (AUTOMATIC1111 vs ComfyUI) with step-by-step instructions.

2. Crafting Effective Prompts for Each Model

No matter the model, a prompt is typically divided into a positive prompt (what you want to see) and a negative prompt (what you don't want). Let's break down best practices for each:

2.1 Positive Prompt Strategies

General Prompt Structure: Most prompts can be thought of as combining a description of the subject with details about style/appearance. A common formula is:

[Subject/scene] with [specific details], [style or medium], [lighting], [quality settings]

For example, "A wizard standing on a misty cliff, holding a glowing staff, digital painting, cinematic lighting, 4K detail". Let's see how to adjust prompt style per model:

Stable Diffusion 1.5 (and its derivatives): Originally, SD1.5 was trained on image captions, but the community found that using terse tags often works better (especially for anime or art styles). For instance, an anime prompt might be: "1girl, blue hair, looking at viewer, masterpiece, best quality, UHD". Here the first part lists subject details (1 girl, hair color, pose) and the latter part lists quality or style tags. Grammatically correct sentences are not necessary – you can string keywords separated by commas. However, for photorealistic outputs you might use more natural phrases like "portrait photo of a woman with soft lighting, 50mm lens, film grain". Experiment with both styles (tag lists vs. descriptive phrases) depending on your model checkpoint (many 1.5-based checkpoints like AnythingV4 or DreamShaper expect tag-style prompts).

Stable Diffusion XL: SDXL was trained with two text encoders and has a larger prompt capacity (up to 2048 tokens internally, though practical UI limits may be lower). It responds best to descriptive prompts written in natural language, followed by style tags. For example: "An astronaut walking on a distant planet, detailed clouds of dust in the air, realistic lighting, 8K photograph, fisheye lens." This mixes a sentence describing the scene with some style keywords. Another example format: A [medium/style] of [subject] [doing something], [additional descriptors]. (tags...). The community often suggests a structure like: "(Style/medium) of (subject) (action/detail). Tags.". For instance: "Anime screencap of a woman with blue eyes wearing a tank top sitting in a bar. Studio Ghibli, masterpiece, pixiv, official art." – notice the sentence followed by comma-separated style tags.

FLUX.1: The guiding principle with FLUX is to "write as if you're talking to a human artist." It parses language in a more sophisticated way, thanks to the T5 transformer. So, you can be quite conversational or specific. For example: "Photorealistic close-up portrait of a medieval knight, intricate engravings on armor, background is a stormy battlefield." You could even write this as multiple sentences or a run-on description – FLUX is robust to that. It doesn't require the "telegraphic" style of comma-separated tags (though you can still provide short phrases if you want; FLUX will handle it). In fact, FLUX excels with detailed, precise descriptions. Feel free to add little details that CLIP models might ignore.

Examples (Positive Prompt for each model):

All three describe a similar scene, but notice the subtle differences in phrasing:

2.2 Negative Prompt Strategies

Negative prompts help you tell the model what to avoid. They're especially useful for removing common artifacts or undesired styles.

Here are best practices by model:

Negative Prompt Example: For a portrait on SDXL, you might use:

Negative prompt: ugly, poorly drawn face, extra fingers, text, watermark, out of frame, low contrast, boring, duplicated

This tells the model to avoid those pitfalls. You can reuse a well-crafted negative prompt across many jobs (some users even save their favorite negative prompt as a text file or use pretrained negative embeddings like "EasyNegative").

Important: Don't include things in the positive prompt that you absolutely don't want, thinking "the model will know I mean the opposite." Models don't understand negation in the positive prompt well. For instance, prompting "a photo of a person without glasses" might actually produce a person with glasses (because it hears "person" and "glasses"). Instead, prompt "a photo of a person" and put glasses in the negative prompt. Always move undesired elements to the negative side explicitly.

2.3 Tips for Specific Styles

Now, some quick tips for achieving popular styles or subjects, and how each model handles them:

Table: Quick Style Keywords

StyleUseful Keywords (add to positive prompt)
Photorealismphotorealistic, DSLR, 35mm, realistic, high detail, 8k, ultra high res, sharp focus, bokeh
Animemasterpiece, best quality, anime illustration, clean line art, flat colors, 2D, character design (plus specific tags for features: e.g. blue hair, school uniform, etc.)
Fantasy Artdigital painting, concept art, epic, fantasy, highly detailed, trending on ArtStation, matte painting, dramatic (and artist names like Greg Rutkowski, John Avon, etc.)
Cinematiccinematic, film still, dramatic lighting, volumetric light, fog, depth of field, motion blur, color graded, 4k
Comic/Cartooncomic book style, ink outline, cel shading, pop art, Pixar style, Disney, cartoon, 2D
Vintage Photoblack and white, vintage, 35mm film, grainy, sepia tone, 1920s, Polaroid, overexposed edges

Use these as inspiration and mix/match. Remember to also adjust your negative prompt if you are aiming for a specific style (e.g., if you want pure anime 2D look, put photorealistic in negative; if you want realistic, put drawing, illustration in negative to avoid cartoonish outputs).

3. Using AUTOMATIC1111 Web UI (Stable Diffusion WebUI)

Now that we've covered what to write in prompts, let's go through how to use the interfaces to generate images. We'll start with AUTOMATIC1111's Web UI (often just called "A1111"), since it's very popular and user-friendly, and then cover ComfyUI in the next section.

Assumption: You have AUTOMATIC1111 WebUI installed with access to the models (SD1.5, SDXL, etc.). If not, follow a Stable Diffusion WebUI installation guide first. Make sure to place your model .ckpt or .safetensors files in the models/Stable-diffusion directory and restart the UI so they appear.

3.1 Loading Models in A1111

Step 1: Launch the Web UI. You'll see the interface with a text area for prompt, one for negative prompt, and options below.

Step 2: Select your model checkpoint. In the top left, there's a drop-down (it might show "Stable Diffusion v1.5" or another model's name). Click it and choose the model you want:

Step 3: VAE (if needed). Some models (especially SD1.5 custom ones) require a VAE (Variational Autoencoder) for color fidelity. A1111 might auto-load a default one. If your outputs have strange colors or contrast, ensure the correct VAE is loaded (in Settings > Stable Diffusion > SD VAE or via the UI's bottom drop-down if visible). SDXL has its own VAE built-in, so usually no action needed there.

3.2 Structuring Prompts and Settings in A1111

Now, type your positive prompt in the text area at the top, and your negative prompt in the bottom text area. Refer to Section 2 for content guidance. Let's go over key settings:

Step 4: Generate the image. Click the Generate button. Wait for the process to complete and your image will appear.

3.3 Using Negative Embeddings (Optional Advanced)

A1111 allows using textual inversion embeddings (tiny .pt or .bin files that represent a concept or style) in prompts. For example, a popular negative embedding is "EasyNegative" which, when put in the negative prompt, can improve general quality of portraits. If you have such an embedding (usually you'd download it and place in embeddings folder), you just type its name in the negative prompt. E.g. negative prompt: EasyNegative, bad-hands-5, (low quality:1.3). The embedding names act like special tokens.

Likewise, there are positive embeddings to invoke styles or specific people. Use these carefully and ensure they match the model version (most embeddings are for SD1.5, they might not work well in SDXL). FLUX likely doesn't support textual inversion (since its text encoder is different).

3.4 Applying LoRAs in A1111

If you want to use a LoRA model to apply a style or character, make sure the LoRA file (.safetensors) is placed in models/Lora. Then:

3.5 ControlNet in A1111

ControlNet is an extension that lets you guide image generation with an input like a pose skeleton, sketch, depth map, or other conditions. For example, you can draw a rough pose stick figure and have SD generate a character in that exact pose.

To use ControlNet (after installing the extension):

ControlNet is extremely powerful for achieving specific compositions, matching a reference, or doing things like inpainting (filling in part of an image), outpainting (expanding an image), etc. Covering all of ControlNet is beyond our scope, but many online guides exist. For now, remember it's a tool to add when you need more control than prompts alone can offer.

3.6 High-Res Fix / Upscaling

If you want a larger final image or more detail:

By now, you should be comfortable using A1111 to generate images with each model. Next, we'll explore ComfyUI, which might seem complex at first but offers powerful control – especially useful for things like SDXL's dual prompts and FLUX.

4. Using ComfyUI for Advanced Workflows (and FLUX)

ComfyUI is a node-based interface for Stable Diffusion. It's extremely flexible: you build a graph of nodes for your pipeline. This makes it perfect for advanced models like SDXL (with dual encoders and refiners) and FLUX (with its custom text encoder), and for doing things like model merges or complex ControlNet setups.

If you're new to ComfyUI, the interface will look like a flowchart editor. You add nodes (each node might do something like "Load Checkpoint", "Text Encode Prompt", "Sampler Step") and connect them. Thankfully, many community workflow files exist that you can load and just edit the prompts.

4.1 Setting Up and Loading a Model in ComfyUI

Step 1: Open ComfyUI. You'll see an empty workspace (grid background). Typically, ComfyUI might come with an example workflow or you can load one.

Step 2: Load a workflow for the model you want. The easiest way to start is to use a pre-made workflow:

Step 3: Identify key nodes to interact with:

Step 4: Enter your prompt in ComfyUI:

Step 5: Set other parameters in the nodes:

Step 6: Run the workflow:

Using ComfyUI effectively: It may be useful to split your workflow into sections:

ComfyUI lets you do model merging as well by using a Model Merge node, but that's beyond basic usage. You can merge two models by loading them and connecting to a merge node with a ratio, producing a new model output. This is an advanced topic; you might use A1111's checkpoint merger for simplicity if you're not comfortable with nodes for merging.

4.2 Special Considerations in ComfyUI for Each Model

4.3 Workflow Example in ComfyUI

Let's walk through an example of generating an SDXL image in ComfyUI to consolidate understanding:

Suppose we want to generate a fantasy castle scene with SDXL and refine it.

  1. Load SDXL base and refiner workflow (for example, the ComfyUI Wiki's SDXL workflow or one shared on Reddit).
  2. In the graph, find the node or section to input prompts. It might have a node labeled "Positive Prompt" and "Negative Prompt" (some workflows create a custom node group for convenience).
  3. Enter: Positive: "A majestic castle on a hilltop overlooking a lake, golden sunset light, high detail, concept art, matte painting, epic atmosphere". Negative: "low quality, oversaturated, cartoonish, people" (we don't want any characters or low quality).
  4. Ensure SDXL base checkpoint is loaded in the CheckpointLoader (e.g. select stable-diffusion-xl-base-1.0.safetensors). The refiner CheckpointLoader should have stable-diffusion-xl-refiner-1.0.safetensors.
  5. Check the resolution: set base generation at 1024×576 (a nice wide aspect for a castle landscape).
  6. Sampler node: use DPM++ 2M Karras, 30 steps, CFG 7 for base. The refiner sampler: maybe 15 steps, CFG 5 (you want a lighter touch on the refiner).
  7. Hit Execute. The base model generates the image. Then the refiner model will run (you'll see the nodes processing sequentially if set up).
  8. Result: An image appears of a castle. If it's too dark or not as expected, adjust prompt or settings and run again.
  9. If satisfied, save the image (ComfyUI usually auto-saves in output folder with a filename). You can also connect an "Save Image" node to auto-save with a custom name if desired.

Now a FLUX example snippet for comparison:

Say we want FLUX to generate the same concept:

  1. Load a FLUX text2img workflow. (We'll assume flux [dev] model loaded.)
  2. Find CLIPTextEncodeFlux node UI. Enter in t5xxl: "A majestic castle on a hilltop overlooking a lake at sunset. The scene is painted in epic fantasy style with golden light and dramatic clouds." In clip_l: "castle, hill, lake, sunset, epic, fantasy".
  3. Negative (depending on workflow, might be a separate node or part of the same): "low quality, blurry, people, text".
  4. Ensure Flux model is loaded in UNet (the loader might have loaded a .safetensors for flux dev).
  5. Resolution: FLUX dev can do 768×432 or bigger; try 768×432 for speed.
  6. Sampler: use Euler or DPM++ SDE, ~25 steps, CFG ~7.
  7. Execute. FLUX will process (taking a bit to load T5). The output appears, likely very coherent with our description.
  8. Compare it with SDXL's output. Perhaps FLUX gave even more vibrant detail or placed things slightly differently because of how it understood the prompt. If needed, refine the wording and re-run.

4.4 Troubleshooting ComfyUI Outputs

5. Advanced Techniques and Tips

Finally, let's cover some advanced techniques that apply to all models and interfaces:

5.1 Using LoRAs for Styles or Subjects

We touched on LoRAs in A1111; in ComfyUI, it's similar conceptually but via nodes. You can load multiple LoRAs at once to mix effects. A common use is adding a LoRA for a specific art style or a specific character face to an existing model.

5.2 Embeddings (Textual Inversion)

Textual Inversion embeddings (like we mentioned "EasyNegative") are custom "words" that the model learns. Many 1.5 embeddings are out there for celebrity faces, styles, etc. They are used by just including the token in prompt, after you load them in the UI.

5.3 Model Merging and Checkpoint Mixes

If you want the flexibility of one model plus style of another, you can try merging:

5.4 Inpainting and Outpainting

Both A1111 and ComfyUI can do inpainting (editing part of an image):

5.5 Prompt Iteration and Evolution

No prompt is perfect first try. A recommended approach:

  1. Start simple – just subject and one style element. Generate a few.
  2. Add details incrementally. See how each affects output. This way you learn each model's behavior.
  3. Use Prompt weighting (for SD1.5/SDXL) or rephrasing (for FLUX) to push the image closer to what you imagine.
  4. If the image is almost there except one thing (say the background is too busy), you can try adding busy background to negative or explicitly say "simple background" in positive.
  5. Keep notes of prompts that work well. You can save prompt text along with images for reference.

5.6 Community Resources

Leverage the community:

5.7 Safety and Ethical Use

A gentle reminder: these models can generate virtually anything – be mindful of not producing disallowed or harmful content. Both A1111 and ComfyUI rely on the user to use them responsibly. Follow the usage guidelines of the model (some models have certain terms that they avoid due to training). Also, credit artists if you heavily use their style, and avoid explicitly copying living artists' works verbatim in prompts if it's against their wishes.


Conclusion and Next Steps

You now have a comprehensive overview of how to get the best results when prompting Stable Diffusion 1.5, SDXL, and FLUX.1, across two powerful interfaces (AUTOMATIC1111 and ComfyUI). To recap:

We included structured breakdowns and comparison tables throughout to illustrate differences. Use this guide as a reference as you create – perhaps keep it open while you work in one window and the UI in another.

Happy prompting, and may your imagination come to life in vibrant detail!

From the studio

Skip the trial and error

Dialing in every setting yourself takes weeks. Digital Divas designs a consistent AI character for you — custom or ready-made — and runs it on our hosted creator platform, with mentors who've done it before.

More guides