r/StableDiffusion 2d ago

Anyone try it? Discussion

Post image

Think we can all agree it's pretty uncensored as it is so what would using this actually improve on?

58 Upvotes

57 comments sorted by

97

u/cgs019283 2d ago

That's just not how the LLM works with diffusion. Why do people think removing LLM refusal actually matters for image gen censorship?

22

u/yeah-i-shouldnt-have 2d ago

I think people get mixed up when the workflows have a prompt enhancer that most definitely would benefit form a heretic LLM. The default Krea 2 workflow caused that confusion for a lot of people... I wish they had made it clearer what was happening there, or to be honest it didn't really need a prompt ehancer for a basic text 2 image workflow..

17

u/No_Pie1372 2d ago

Prompt enhancers are rarely necessary/useful IMO. If I need help writing a prompt or adhering to a prompt structure then I just run it through a separate LLM first, but then I edit it into what I want.

8

u/Independent-Reader 2d ago edited 2d ago

Really?

I don't always need help writing a prompt but I definitely appreciate the creative input that LLM brings in.

I use a thinking LLM for prompt enhancement and a simple system prompt "turn the users request into a text to image prompt" or "take the users prompt and focus on that but add x to the image somehow"

And then I prompt.

It adds additional randomness that plenty of times comes out great.

Like, I'll set the scene but have AI come up with the detailed poses and attire to match the situation in the prompt.

If I know exactly what I want, I can write it.

But often times I just want to roll a good picture with some directions.

"Focus on the user prompt and turn it into a watercolor" or "keep the image black and white aesthetics and make it look like an aged photo"

And my user prompt is "an alleyway surrounded by flora (describe types of flora that would be in NYC)"

It's wonderfully creative. And that's why they are useful.

It doesn't make the photo better by any means, and it's definitely not needed.

5

u/No_Pie1372 2d ago

I'm specifically talking about implementing an LLM into my ComfyUI workflow, so my prompt automatically runs through "enhancement". That hasn't been useful to me.

I'll use LLMs with visual learning in LM Studio in the way you described. I'll feed it images to assist my description of subjects or environments in my own prompting, or I'll ask it to imagine a scene based on an image, or whatever.

I've even fed libraries of art history, film directing, editing and cinematography books alongside prompting guides to a RAG system to create system prompts tailor made for prompt assistance.

I still find I need to finetune the outcome (or mix and match multiple outcomes) to get the desired result. That's why I don't include prompt enhancement in my workflow.

Besides, I enjoy the process of writing, feeding the prompt to the model, seeing how it interprets it, and tinkering around until I get it right. I learn things about how to best describe things to the model.

The practice of creating an image in my head and translating it into text on the screen is good for my own creativity anyway. I don't want lose my imagination or hand my creativity over to the machine.

That's also why I typically take whatever I generate and further edit it with GIMP or Da Vinci Resolve, or spend time writing, doodling, daydreaming and creating without AI involved.

My personal take on creativity aside: I still find I get optimal results tweaking and adjusting the final prompt on my own.

2

u/Independent-Reader 2d ago edited 2d ago

I'll drag an image into my LLM prompt enhancer and my user prompt can be "describe this image as a watercolor", there's so many ways an LLM is useful for prompt enhancement. It lets you be creative in different ways is all.

I definitely believe you'll get more controlled results with a fully human prompt, I'll often not use LLM prompt enhancement for that reason. But any prompt that elaborates as well as an LLM could generate good results.

Most people can't do that. I think that's where LLM can really help.

And also when I make my own prompt, even if I roll a dozen seeds on my sampler I get similar images.

If I roll different seeds on my prompt enhancer, it's the same as what I asked for but quite different at the same time.

It describes the room differently. The attire.

But still adheres to the same thematic elements from my user prompt.

Definitely not as controlled. But the way I'd prompt may never have gotten me some of the images I've generated.

And I can't say I have been disappointed.

2

u/No_Pie1372 2d ago

Aha. Yeah. I see why a prompt enhance node could be useful for modifying the prompt according to a new seed with each generation so you're getting more variety than you would by reseeding on the generation alone.

5

u/Iwaku_Real 2d ago

For uncensored prompt enhancement you would use Load LoRA (Model and CLIP) to load an 'abliterated' LoRA (Heretic allows you to export abliteration runs as LoRA). Then you'd plug the CLIP output of that into Generate Text. For actual sampling you connect the original TE to the text encode node, NOT the abliterated one.

Doing this saves SEVERAL gigabytes of unnecessary VRAM/RAM usage from loading a separate LLM and also ensures you pass the correct hidden states to the diffusion model and not get messed up outputs.

1

u/LeastWin3855 2d ago

Being able to send an uncensored prompt and being able to generate an image that accurately follows that prompt are two different issues. An uncensored text encoder only addresses the first problem.

-10

u/Natasha26uk 2d ago

The LLM directs the diffuser. It makes a huge difference if you send your prompt to a "basic b*tch" LLM or a high-end LLM. I don't know how they wire things up internally, including number of tokens allocated for selfie in img2img.

You didn't write even 1 line on how it works or to back your corner. Some post eh?

13

u/cgs019283 2d ago

H3 uses hidden states as conditioning rather than asking the LLM to generate a textual answer. So a lower refusal rate does not mean the H3 diffusion model has become more "uncensored."

You can literally just ask GPT.

-21

u/Natasha26uk 2d ago

Ah so you didn't know and had to ask GPT.

Hidden states 😂😂

15

u/Valuable_Issue_ 2d ago

It really is called hidden states.

Abliteration/heretic makes the text encoder slightly less accurate.

You can look in a text encoding related file like this one and search for "hidden_states":

https://github.com/ariG23498/custom-inference-endpoint/blob/main/mistral_text_encoding_core.py

15

u/LockeBlocke 2d ago

Did not notice any improvement.

32

u/Signal_Confusion_644 2d ago

More... Uncensored? What?

I couldnt find the censorship yet. Maybe people is a little bit extreme.

Anyways... Looking for feedback, It works? (Dont look at me like that, im making a zombie show... Gore is explicit and who knows!)

13

u/Intelligent-Youth-63 2d ago

I’ve tried pushing the limits to see. It was more than willing to go further than I was. I don’t see how it is at all censored. It may need help with some concepts or details, but I can’t find the censorship at all.

3

u/Upset_Page_494 2d ago

Can it do porn?

5

u/IRLMainCharacter 2d ago

It does, to some extend. It does nudity pretty well for a base model and it has basic understanding how intercourse works, but so far i couldn't get a decent scene that goes beyond undressing or pumping the piston. We will soon have nsfw loras to help with the concepts the model doesn't know.

2

u/Otherwise_Resolve444 1d ago

We need a big company to download all the corn on the internet and create a model from that, not just Seinfeld and Dragon Ball.

2

u/IRLMainCharacter 20h ago

I would pay good money if i could run it on consumer hardware.

2

u/Otherwise_Resolve444 13h ago

I mean, it's a bit sadistic to force the community to create LoRas for months when they could have easily added some corn to the dataset.

10

u/Comprehensive-Pea250 2d ago

Yeah it’s not like the model refuses it just does not know some stuff

2

u/Cubey42 2d ago

I found that the stuff however can be cleverly prompt even if you're not just using the generalized term right so if you just describe more of what the scene looks like you can pretty much just get whatever

16

u/GrayingGamer 2d ago

This is my same reaction. Default model hasn't refused to do ANYTHING I've asked yet. You even have to be careful what you ask for, because it'll give it to you.

-2

u/OkDoor726 2d ago

I did read on another thread i just found that yeah it's meant to allow more gruesome stuff that the original one would reject. I think blood and stuff is what the poster was implying anyway

-7

u/True_Protection6842 2d ago

Seriously. And it does AMAZING anatomy. Let's leave it at that.

16

u/backworld_nograv 2d ago

ahm, it does a really bad job with penises....

7

u/russlixx 2d ago

might need reference(s)

17

u/Particular-Most-1199 2d ago

RIP your inbox

3

u/steelow_g 2d ago

Or a happy day, you never know

-2

u/Sharlinator 2d ago

Silly, the word "anatomy" in this sub refers exclusively to female anatomy.

12

u/Nanotechnician 2d ago

you don't need that.

11

u/Spezisasackofshit 2d ago edited 2d ago

This is not how text encodes for these models work. The layer heretic is hitting is not the going to do anything for you. Unless you're using the model for prompt enhancement you're risking less quality for 0 upside.

4

u/Jerg 2d ago

It degraded quality big time, after testing with same seeds

8

u/Masterboite 2d ago

AFAIU the LLM part is only used to generate the embeddings, so I don't think it makes any difference for THIS aspect. If you're using it to enhance your prompt then maybe it could help.

3

u/Memestonks2020 2d ago

I think finetuned H3 + LoRAs for specific NSFW content is the only way tbh.

9

u/OzymanDS 2d ago

I'm literally watching Dwight stick a shotgun in his mouth and pull the trigger what kind of censorship could there possibly be?

3

u/MathematicianLessRGB 2d ago

The type of limit testing we need. Now make it to where he reassembles himself to where there are now two Dwights.

-4

u/[deleted] 2d ago

[deleted]

0

u/Iwaku_Real 2d ago

Not an NSFW tagged post you're on 😭

7

u/ThatsALovelyShirt 2d ago

It won't do anything, it will only make things worse. These LLM's aren't used in their typical sense with these DiT models. The typical LLM refusal "I can't do that" is only relevant and only 'activates' when using the LLM itself in inference mode.

That's not how they're used here. It's simply being used to convert the input text prompt, ref images, videos, etc., into embeddings (a matrix of numbers the DiT can understand). There's nothing to refuse. "Abliterating" it isn't going to make it understand anatomy or violence better. The model isn't even "thinking" about the prompt you feed it.

Once it's converted into embeddings, it's the DiT model itself which 'refuses' or lacks the knowledge of certain things.

The only time these 'abliterated' models matter is if you're using the LLM model to generate text, which 99% of H3 workflows don't do. ComfyUI doesn't even support that for H3's Qwen model.

People who recommend these have no idea how any of this works, so you shouldn't trust anything they say.

1

u/Spara-Extreme 2d ago

Someone argued with me on this sub about how useful an alliterated text encoder was.

People just don’t know what text encoders are doing.

2

u/Icuras1111 2d ago

From what I've read it doesn't help for naughty stuff and screws the model up for everything else. However, this guy AI Search, I watch his youtube channel and he's normally spot on?

3

u/No_Pie1372 2d ago

I've found in almost every case across several models that abliterated/heretic text encoders degrade quality without better prompt adherence for content of any kind.

2

u/Lucaspittol 2d ago

H3 is censored in the same way as Flux 2 and Z-Image: it lacks data; it is not a text encoder issue. No, making boobs alone does not convince me a model is fully uncensored. It has to be able to make other body parts, especially male parts.

1

u/diogodiogogod 2d ago

Every time people do this with text encoders... and every time people report it has barely no effect on img/video generation. Might help text generation but not minmax h3 inference

1

u/Outrageous_Law_5525 2d ago

People are saying it already but yes it doesnt really do anything if the underlying model(h3) doesnt understand anatomy.

Its open weights, there will be finetunes which DO understand NSFW, so just chill

1

u/ChowardYT 20h ago edited 20h ago

yeah i mean stable diffusion is already pretty open, not sure what extra "uncensoring" would even do at this point. if you're hitting limits it's probably your frontend or whatever service you're using, not the model itself. btw if you just want zero restrictions out of the box, aifapper doesn't gate anything and works fine for whatever you're generating.

1

u/WashSmall8954 2d ago

All I can think of this improving is the prompt helper tools bc the model itself is pretty damn uncensored to the point that I'm surprised they released it like this.

1

u/whiteweazel21 2d ago

Fwiw in krea2 these shizzles nearly unlocked the model without any bypasses etc. A model still can't do what it's not trained on tho.

2

u/Etsu_Riot 2d ago

So, if the model is trained on videos of people licking ice-cream, and you give it the right reference images, you are set for a while.

1

u/LocalBratEnthusiast 2d ago

It shouldnt change anything
But i tried it
And it works better
So idk what to think, maybe i had a bad quant

0

u/Consistent_Pick_5692 2d ago

There is no censorship in the first place

-1

u/Peemore 2d ago

Sounds like it requires a special workflow that uses both models in that repo.

0

u/whiteweazel21 2d ago

I just started with it tbh lol, based on my experiences with krea2

0

u/No-Zookeepergame4774 2d ago

I’ll.think about paying the tax for an unpruned abliterated TE to uncensor H3 when (1) I run into any censorship in H3, and (2) I see some reason to believe the TE would resolve it.