I also think that Gemini is usually better than Claude. It seems to be less shy and just as smart.
I really like GLM 5, fast and still very smart. I calibrate most of my scenarios based on what GLM can achieve.
I use Claude for very complex scenarios with very long context sizes, and I think Claude is the best in this regard.
DeepSeek V4 is a disappointment. Its quality seems to be extremely variable depending on the time of day.
Other models can still be pretty good depending on the use case and your scenario/system prompt.
For tone, I recommend adding small examples (in the persona's description) of how the character should talk, language quirks, register, etc.
Use the persona examples only when you need the LLM to have examples of specific response structures: when you want to enforce usage of certain modules, stims, etc.
Last thing, this can be a bit immersion-breaking, but it's an extremely powerful tool: during a session, use your own [OOC message] to guide/fix the LLM.
If Claude isn't stopping stims correctly (it's true that many models prefer to change the mode to tickles + 1 intensity instead of stopping), you can, during the session or even at the very start, add a short OOC message as a header to your message:
[OOC Message from the player to the LLM. Thank you for embodying {{char}} and playing with me! One important thing: please make sure to completely stop the stims, to...].
OOC messages are usually really powerful. If the model is self-censoring, ask it directly not to shy away, to do crude, explicit content, etc.
Your scenario is what drives the experience. Don't hesitate to edit your scenarios and write structured approaches to fix the behaviors you dislike.
For example, I feel that LLMs don't handle edging correctly most of the time. Don't hesitate to give the model a structure to follow in your scenario:
<edge_flow>
When {{char}} makes the user edge, always use the <Delay /> module to make {{user}} hold for a duration. Follow the structure below:
- Short command such as: "Edge, slave!"
- Use a <Delay /> command as the duration of the edge.
- Completely stop the stim at the end of the edge.
Note: You can use multiple edges in a row in a single AI message, with several <Delay />.
</edge_flow>
Same for the stimulation, don't hesitate to add sections, small examples, guidelines in your scenario to explain how to use the stimulation?