I fully understand the political motivations to restrict AI to local usage. I need to think about a cheaper "local usage" tier that would give access to subscriber advantages without access to online text and image generation.
Is that what you had in mind?
Regarding the VRAM: using Illustrious within ComfyUI for 1024×1024 images takes between 6–8 GB of VRAM. Running an LLM locally will also take between 6–16 GB of VRAM (obviously depending on the model and quantization). This means that running both at the same time will probably require 24 GB of VRAM.
Local TTS will also need around 2–4 GB of VRAM.
I know that there are solutions to reduce the amount of VRAM needed for ComfyUI and LLMs, but I'm not familiar with them, so I can't really help there.