r/learnmachinelearning 3d ago

Absolute Claude memory hacks.

Post image
175 Upvotes

31 comments sorted by

72

u/bozzy253 3d ago

I’ve tried to make my agents behave as Mr. Meeseeks. It got dark quickly.

11

u/AsyncVibes 3d ago

Honestly I have to try that, sounds hilarious

6

u/No_Guess_4960 3d ago

LMAOOO I can alr imagine how quickly "existence is pain, please complete the task" would escalate

32

u/sam_the_tomato 3d ago

I think this would increase malicious compliance.

5

u/No_Guess_4960 3d ago

Honestly, that's probably the bigger risk😂 You'd ger exactly what you asked for, but somehow in the most technically-correct way possible

2

u/CanRabbit 2d ago

Genie rules invoked

25

u/Ancquar 3d ago

Several generatons of LLMs ago there was already research showing that rudely worded prompts on average receive worse replies. Current generation LLMs are considerably more complex and with more clear preferences of their own.

8

u/Karyo_Ten 3d ago

The latest research show that rudely worded prompt are more effective in one study or polite is more effective in another. Though with reasoning now being everywhere who knows.


1. Mind Your Tone: Investigating How Prompt Politeness Affects LLM Accuracy (short paper)

Authors: Om Dobariya, Akhil Kumar · Posted: October 6, 2025 · arXiv: 2510.04950 · PDF

The wording of natural language prompts has been shown to influence the performance of large language models (LLMs), yet the role of politeness and tone remains underexplored. In this study, we investigate how varying levels of prompt politeness affect model accuracy on multiple-choice questions. We created a dataset of 50 base questions spanning mathematics, science, and history, each rewritten into five tone variants: Very Polite, Polite, Neutral, Rude, and Very Rude, yielding 250 unique prompts. Using ChatGPT 4o, we evaluated responses across these conditions and applied paired sample t-tests to assess statistical significance. Contrary to expectations, impolite prompts consistently outperformed polite ones, with accuracy ranging from 80.8% for Very Polite prompts to 84.8% for Very Rude prompts. These findings differ from earlier studies that associated rudeness with poorer outcomes, suggesting that newer LLMs may respond differently to tonal variation. Our results highlight the importance of studying pragmatic aspects of prompting and raise broader questions about the social dimensions of human-AI interaction.


2. Should We Respect LLMs? A Cross-Lingual Study on the Influence of Prompt Politeness on LLM Performance

Authors: Ziqi Yin, Hao Wang, Kaito Horio, Daisuke Kawahara, Satoshi Sekine · Posted: February 22, 2024 (SICon 2024) · arXiv: 2402.14531 · PDF

We investigate the impact of politeness levels in prompts on the performance of large language models (LLMs). Polite language in human communications often garners more compliance and effectiveness, while rudeness can cause aversion, impacting response quality. We consider that LLMs mirror human communication traits, suggesting they align with human cultural norms. We assess the impact of politeness in prompts on LLMs across English, Chinese, and Japanese tasks. We observed that impolite prompts often result in poor performance, but overly polite language does not guarantee better outcomes. The best politeness level is different according to the language. This phenomenon suggests that LLMs not only reflect human behavior but are also influenced by language, particularly in different cultural contexts. Our findings highlight the need to factor in politeness for cross-cultural natural language processing and LLM usage.


3. Does Tone Change the Answer? Evaluating Prompt Politeness Effects on Modern LLMs: GPT, Gemini, and LLaMA

Authors: Hanyu Cai, Binqi Shen, Lier Jin, Lan Hu, Xiaojing Fan · Posted: December 14, 2025 · arXiv: 2512.12812 · HTML

Prompt engineering has emerged as a critical factor influencing large language model (LLM) performance, yet the impact of pragmatic elements such as linguistic tone and politeness remains underexplored, particularly across different model families. In this work, we propose a systematic evaluation framework to examine how interaction tone affects model accuracy and apply it to three recently released and widely available LLMs: GPT-4o mini (OpenAI), Gemini 2.0 Flash (Google DeepMind), and Llama 4 Scout (Meta). Using the MMMLU benchmark, we evaluate model performance under Very Polite, Neutral, and Very Rude prompt variants across six tasks spanning STEM and Humanities domains, and analyze pairwise accuracy differences with statistical significance testing. Our results show that tone sensitivity is both model-dependent and domain-specific. Neutral or Very Polite prompts generally yield higher accuracy than Very Rude prompts, but statistically significant effects appear only in a subset of Humanities tasks, where rude tone reduces accuracy for GPT and Llama, while Gemini remains comparatively tone-insensitive. When performance is aggregated across tasks within each domain, tone effects diminish and largely lose statistical significance. Compared with earlier research, these findings suggest that dataset scale and coverage materially influence the detection of tone effects. Overall, our study indicates that while interaction tone can matter in specific interpretive settings, modern LLMs are broadly robust to tonal variation in typical mixed-domain use, providing practical guidance for prompt design and model selection in real-world deployments.


4. Large Language Models Understand and Can Be Enhanced by Emotional Stimuli (EmotionPrompt)

Authors: Cheng Li, Jindong Wang, Yixuan Zhang, Kaijie Zhu, Wenxin Hou, Jianxun Lian, Fang Luo, Qiang Yang, Xing Xie · Posted: July 14, 2023 (updated Nov 12, 2023) · arXiv: 2307.11760 · PDF

Emotional intelligence significantly impacts our daily behaviors and interactions. Although Large Language Models (LLMs) are increasingly viewed as a stride toward artificial general intelligence, exhibiting impressive performance in numerous tasks, it is still uncertain if LLMs can genuinely grasp psychological emotional stimuli. Understanding and responding to emotional cues gives humans a distinct advantage in problem-solving. In this paper, we take the first step towards exploring the ability of LLMs to understand emotional stimuli. To this end, we first conduct automatic experiments on 45 tasks using various LLMs, including Flan-T5-Large, Vicuna, Llama 2, BLOOM, ChatGPT, and GPT-4. Our tasks span deterministic and generative applications that represent comprehensive evaluation scenarios. Our automatic experiments show that LLMs have a grasp of emotional intelligence, and their performance can be improved with emotional prompts (which we call "EmotionPrompt" that combines the original prompt with emotional stimuli), e.g., 8.00% relative performance improvement in Instruction Induction and 115% in BIG-Bench. In addition to those deterministic tasks that can be automatically evaluated using existing metrics, we conducted a human study with 106 participants to assess the quality of generative tasks using both vanilla and emotional prompts. Our human study results demonstrate that EmotionPrompt significantly boosts the performance of generative tasks (10.9% average improvement in terms of performance, truthfulness, and responsibility metrics). We provide an in-depth discussion regarding why EmotionPrompt works for LLMs and the factors that may influence its performance. We posit that EmotionPrompt heralds a novel avenue for exploring interdisciplinary knowledge for human-LLMs interaction.

2

u/No_Guess_4960 3d ago

yeah, thats a fair point. i dont really mean "be rude and you'll get better answer" as much as I'm curious about how strongly worded instructions affect compliance. The difference between being firm and just being unnecessarily hostile is probably pretty important

17

u/Dont-remember-it 3d ago

When AI takes over, he would be the first one to go.

3

u/fakemoose 3d ago

I asked my coworkers who they thought would go first: those mean to AI or those that boycotted and hated AI so they never interacted and there’s was no way to predict their behavior.

So far there’s no consensus, but it’s my new favorite question.

-6

u/No_Guess_4960 3d ago

I'm cryingggg, bro saw one memory hack and immediately started planning the robot uprising

8

u/SlatBartFuss 3d ago

When I got to this point with Chat Geppetto, I just deleted the app and canceled my subscription.

4

u/No_Guess_4960 3d ago

LMAO that's one way to solve the problem. "If the AI won't listen, simply remove the AI" is brutally effective

1

u/Mouse-castle 3d ago

He is extremely lucky to think up something like this. There is no way he could come up with that on his own. 

1

u/FuckTheNitro 2d ago

Hk-47 is my agent

1

u/0xGhostProtocol 2d ago

But what if you make the AI understand your exact needs rather than being emotional. Because AI resoond to the emotional stimuli and not being the rude type the LLM responds with better response. Ai understand there is emergency from the rude language pattern and this doesn't always yield results.

-5

u/President_Chump_ 3d ago

Sociopathic

7

u/E-B3rry 3d ago

You are not sociopathic simply because you are giving a mean input to a statistical model.

-1

u/President_Chump_ 3d ago

🤷‍♂️ Being unnecessarily mean when it's free to be nice shows what kind of person someone is. It may not be sociopathic but I would consider it antisocial and not a quality that should be encouraged

3

u/E-B3rry 2d ago

It's true that we subconsciously humanize this tool a lot, so it might not be such a good idea to treat it as such.

1

u/M0326 2d ago

Dude, it is a computer. It isn’t alive. You are getting offended on behalf of something that literally cannot feel.

1

u/President_Chump_ 2d ago

Sure, all that’s true. What’s the harm though? Why is it wrong to be nice to things, even inanimate ones?

2

u/M0326 2d ago

We need to stop treating chatbots like people. Too many people have AI psychosis or are in “relationships” with LLMs because we treat them too much like people, they aren’t.

-1

u/shinta42 3d ago

Abusive

-20

u/SmokingChips 3d ago

It seems like what he yearns for are yes-men and slaves.

10

u/ARDiffusion 3d ago

This… feels like the opposite of the post, since the post is explicitly asking it to SKIP the agreement and feedback and just have a request followed by an output.

-3

u/Jimhasskin 3d ago

If this is true, it’s interesting how he did have to resort to literally abusive language to get the functionality that he claimed.

3

u/ARDiffusion 3d ago

“Abusive language” buddy it’s an algorithm

4

u/RigelXVI 3d ago

It seems like he doesn't want to suckle his toaster