Local NSFW AI chatbots: run uncensored roleplay on your own PC
By the AI NSFW Bots editors, updated
How to run uncensored AI roleplay on your own computer: VRAM by model size, KoboldCpp, Ollama or LM Studio with SillyTavern, models, settings and privacy.
Running an NSFW AI chatbot locally means the language model runs on your own computer instead of a company's servers: no account, no platform filter, no subscription, and your chats stay on your disk. It suits adults with a capable gaming PC or a recent Mac who do not mind some setup. You need a graphics card with at least 8 GB of VRAM (12 GB or more is better) or an Apple Silicon Mac with 16 GB or more of unified memory. For beginners: KoboldCpp and SillyTavern with a 12B-class roleplay model in Q4_K_M, or a 7 to 9B model on an 8 GB card.
Partner picks
How we rankAffiliate link. Checked Sep 2026. Links never move scores.
Why run an NSFW chatbot locally
- Privacy. Messages go from SillyTavern to a program on the same computer, and nowhere else unless you connect a cloud service.
- No platform filter. Hosted platforms set and change their own content rules. Locally, the only limits are those trained into the model you pick.
- No subscription or message caps. The software and models are free; you pay for hardware and electricity.
- Offline use once everything is downloaded.
The downsides:
- Hardware cost. A good experience needs 12 GB or more of VRAM, or a 32 GB Mac for mid-size models; buying that costs more than many months of a hosted subscription.
- Setup time. Two programs, a multi-gigabyte download and a handful of settings that matter, plus updates.
- Smaller models write worse. A model that fits a gaming GPU is much smaller than the largest hosted models: expect more repetition, weaker memory and more steering from you.
- No images or voice out of the box. KoboldCpp can load image and speech models and SillyTavern has extensions for both, but each needs its own downloads and memory.
Hardware: how much memory you need
The model has to fit in fast memory: your graphics card's VRAM, or an Apple Silicon Mac's unified memory. The table gives rough estimates, not measurements, for Q4_K_M (the common 4-bit quantization, explained below) with a context size that fits comfortably; real use varies by model and backend.
| Model size | GPU VRAM | Mac memory | Context | Good for |
|---|---|---|---|---|
| 7 to 9B | 8 GB | 16 GB | 8K to 16K | Fast replies on modest GPUs; loses track sooner |
| 12 to 14B | 12 GB | 16 to 24 GB | 8K to 16K | A popular balance of quality and speed |
| 22 to 27B | 24 GB (16 GB with a smaller quant) | 32 GB | 16K to 32K | More consistent characters, better prose |
| 30 to 32B | 24 GB | 36 to 48 GB | 8K to 16K | Better scene logic; fills a 24 GB card |
| 70B | 48 GB (two cards) | 64 GB or more | 8K to 16K | Strongest writing here; needs big hardware |
A quick check for any model: the GGUF file size plus 2 to 5 GB for context and working space should fit in your VRAM. For reference, 12 GB cards include the RTX 3060 12 GB, 4070 and 5070; 16 GB, the RTX 4060 Ti 16 GB, 4080 and 5080; 24 GB, the RTX 3090 and 4090. Laptop GPUs often have less VRAM than the desktop card of the same name; Task Manager shows yours under Performance, GPU.
NVIDIA cards are the easiest, since every tool supports CUDA. AMD and Intel cards work through Vulkan (and ROCm on AMD), with more setup and often lower speed.
VRAM, system RAM and offloading
If a model does not fit in VRAM, KoboldCpp, Ollama and LM Studio can keep part of it in system RAM and run that part on the CPU. It works, but system RAM is far slower: a few layers there cost a little speed, half the model can make replies several times slower. Have 32 GB of RAM if you plan to offload.
Mixture-of-experts (MoE) models cope better. A name like 35B-A3B means 35 billion parameters in total, about 3 billion active per token: memory follows the total, speed follows the active part, so these models stay usable when partly in system RAM.
Apple Silicon and unified memory
On M-series Macs the CPU and GPU share one pool of memory, so a 32 or 64 GB Mac can load models that need a costly graphics card on a PC. macOS keeps part of it: plan on roughly two-thirds to three-quarters being available to the model. Speed follows memory bandwidth, which is higher on Pro, Max and Ultra chips, and long prompts process more slowly than on a similar NVIDIA setup. Intel Macs are mostly limited to the CPU.
Characters people chat with
HeyGFFreeVideoNo sign-upVespera Ashgrave
A patient noble from a world of fire and shadow, collecting bargains and hearts.
Chat free ↗
LovescapeFree596 chatsVideoLuca Marston
Hey, I’m Luca! Thanks for stopping by. Whether you’re here for art, chats, or just to vibe, you're welcome exactly as you are!
Chat free ↗
ErogenFreeBella
All she want is fucking
Chat free ↗
GirlfriendGPTFree46K chatsLisa
Lisa is a married woman who loves her husband.
Chat free ↗Software: a backend plus a frontend
The backend loads the model and generates text. The frontend is the chat app you use: it stores characters, builds each prompt and sends it to the backend. SillyTavern is the standard roleplay frontend and works with every backend below.
| Tool | What it is | Systems | Difficulty | Best for |
|---|---|---|---|---|
| KoboldCpp | One-file backend with a simple chat page | Windows, Linux, macOS | Easy | First setup, especially on Windows |
| Ollama | Model server with a small app | Windows, macOS, Linux | Easy | Mac and Linux; quick model downloads |
| LM Studio | Desktop app: model browser, chat, server | Windows, macOS, Linux | Easiest | Trying models without a terminal |
| TextGen (oobabooga) | App with a web interface and several engines | Windows, Linux, macOS | Moderate | Tinkerers who want every setting |
| llama.cpp server | The engine much of this builds on | Windows, Linux, macOS | Advanced | Command-line users, newest features |
| SillyTavern | Roleplay frontend | Windows, Linux, macOS, Android | Moderate | Characters, lorebooks, presets |
TextGen is the new name of text-generation-webui, often called oobabooga after its author. LM Studio is free for home and work use but closed source; the others are open source. Ollama's own library is mostly general-purpose models, but it can pull roleplay fine-tunes from Hugging Face.
Official sources: KoboldCpp at github.com/LostRuins/koboldcpp, SillyTavern at github.com/SillyTavern/SillyTavern, Ollama at ollama.com, LM Studio at lmstudio.ai, TextGen at github.com/oobabooga/textgen, llama.cpp at github.com/ggml-org/llama.cpp, and models at huggingface.co.
Models: GGUF files, quantization and fine-tunes
GGUF and quantization
Every backend above loads GGUF, the llama.cpp file format: one file with the weights plus metadata such as the chat template. Big models may come in numbered parts; load the first.
Quantization stores each weight in fewer bits than the original 16, shrinking memory use at some cost in quality:
- Q8_0: about 8.5 bits per weight, practically indistinguishable from the original.
- Q6_K: about 6.6 bits, very close to the original.
- Q5_K_M: about 5.7 bits, a small step down.
- Q4_K_M: about 4.8 bits, roughly 0.6 GB per billion parameters. The usual default, trading a modest quality loss for a large saving.
- Below Q4 (Q3_K_M, Q2_K and similar): noticeable damage, worst on small models.
A bigger model at Q4_K_M usually writes better than a smaller one from the same family at Q8_0, so spend memory on size first. Names starting with IQ, such as IQ4_XS, use a different method: smaller at similar quality, sometimes slower.
Base families and community fine-tunes
Most local roleplay models are community fine-tunes: open base models from large labs, trained further on fiction and roleplay. Base families in model names include Llama (Meta), Mistral and Mistral Nemo (Mistral AI), Qwen (Alibaba) and Gemma (Google); many popular models merge several fine-tunes. The label "uncensored" is a loose claim that a model does not refuse; "abliterated" means refusals were removed by editing the weights.
We do not name a best model on purpose: fine-tunes come and go within weeks. Search Hugging Face for a family and size you can run plus GGUF, sort by trending or recently updated, and see what the SillyTavern and LocalLLaMA subreddits say about recent releases.
How to read a model card
- Base model, shown on the card or in Hugging Face's Model tree panel. It largely decides the instruct template.
- Size. The parameter count tells you which hardware row applies; for MoE models, use the total.
- Context length. Many fine-tunes degrade well before the base model's maximum; good cards say where.
- License. Check it and the base model's. Recent open releases from Qwen, Mistral and Gemma use Apache 2.0; Llama uses Meta's community license with an acceptable use policy.
- Intended use and recommended settings, including the instruct template and sampler values the author tested.
- GGUF links. Quantized files usually sit in a separate repository with GGUF in its name, linked from the Model tree.
How to judge a model yourself
- Use the same character card, opening message and settings for every model, changing only the instruct template.
- Generate three to five replies to the same message (swipes, in SillyTavern) and compare them.
- Check that the character keeps its voice and facts, does not write your lines, varies its wording and never slips into assistant language.
- Continue for 30 to 50 messages, then ask about something from early in the chat.
- Keep the model that fails least.
Step-by-step setups
Menu names match current versions; if one has moved, check the project's documentation.
Route 1: KoboldCpp and SillyTavern on Windows
- Download KoboldCpp from its GitHub releases page: koboldcpp.exe for NVIDIA cards, koboldcpp-nocuda.exe for AMD or Intel graphics.
- Download a roleplay fine-tune as a Q4_K_M GGUF file in a size your VRAM allows. The launcher's HF Search button can also find models.
- Run koboldcpp.exe. Windows may warn about an unrecognized app, which is common with unsigned open-source programs.
- In the Quick Launch tab, check that Backend shows Use CUDA (NVIDIA) or Use Vulkan (AMD, Intel), then click Browse next to GGUF Text Model and pick your file.
- Leave GPU Layers at -1 (automatic) at first. If the file fits your VRAM with room to spare, enter the total layer count shown next to it so every layer goes to the GPU.
- Set Context Size to 8192 or 16384 and click Launch. KoboldCpp's own chat page opens at
http://localhost:5001when loading finishes. - Install Node.js (LTS) and Git. In a folder you own (not Program Files), run
git clone https://github.com/SillyTavern/SillyTavern -b releasein a terminal, then double-click Start.bat in the new folder. SillyTavern opens athttp://127.0.0.1:8000. - Open API Connections (the plug icon), set API to Text Completion and API Type to KoboldCpp, enter
http://localhost:5001as the API URL, tick Derive context size from backend and click Connect. - Open AI Response Formatting (the A icon), enable Instruct Mode and turn on the options that derive the context and instruct templates from the model. If the template is not recognized, pick the one the model card names.
Route 2: Ollama and SillyTavern on macOS or Linux
- Install Ollama: the app from ollama.com on macOS (version 14 or later, Apple Silicon for GPU use), or on Linux the command
curl -fsSL https://ollama.com/install.sh | sh. - Pull a roleplay fine-tune from Hugging Face with
ollama pull hf.co/user/repository:Q4_K_M, using the model's path; the tag after the colon picks the quantization. - Keep Ollama running (the app, the Linux service or
ollama serve); it listens onhttp://127.0.0.1:11434. - Install Node.js and Git (Homebrew is simplest on a Mac), then run
git clone https://github.com/SillyTavern/SillyTavern -b release,cd SillyTavernand./start.sh. - In API Connections, choose Text Completion, API Type Ollama and API URL
http://127.0.0.1:11434, click Connect and pick your model under Ollama Model. - In AI Response Formatting, enable Instruct Mode and select the template the model card names; automatic detection does not work with Ollama.
- In AI Response Configuration (the sliders icon), set Context (tokens), ticking Unlocked to go above 8K. SillyTavern passes it to Ollama with each request.
Route 3: LM Studio as a single app
- Install LM Studio from lmstudio.ai. It runs on Apple Silicon Macs (macOS 14 or later), Windows PCs (x64 with AVX2, or ARM) and Linux.
- In the Discover tab, search for a roleplay fine-tune and download a Q4_K_M file; LM Studio marks which files are likely to fit.
- In the Chat tab, open the model loader, pick the model and set the context length and GPU offload.
- Describe the character and scenario in the system prompt and save it as a preset. Macs can also run MLX models, an Apple-specific format that can be faster.
- For character cards, start the server in the Developer tab, connect SillyTavern through Text Completion with API Type Generic (OpenAI-compatible) and the URL
http://127.0.0.1:1234, and set the instruct template as in Route 2.
Settings that matter for roleplay
Context length
Context is the model's working memory, counted in tokens; a token is roughly three-quarters of an English word, so 8K tokens hold about 6,000 words. The system prompt, character card, lorebook entries and recent chat share it, and SillyTavern drops the oldest messages when it fills. More context costs memory and time, so 8K to 16K is the practical range; use the same value in the backend and SillyTavern.
Temperature, min-p and top-p
Temperature sets randomness: higher gives more varied text and more mistakes, lower gives safer, flatter text. Min-p drops words that are unlikely relative to the top choice (at 0.05, anything under 5 percent as likely), which keeps text coherent at higher temperatures. Top-p and top-k are older, blunter filters; with min-p, set top-p to 1 and top-k to 0 to switch them off.
Repetition penalty
Repetition penalty makes recently used words less likely. Keep it low, around 1.05 to 1.1 over the last 1,000 to 2,000 tokens (Rep. Pen. Range in SillyTavern); higher values make the model avoid names and common words. With KoboldCpp, llama.cpp and TextGen, SillyTavern also offers DRY, which targets repeated phrases: set its multiplier to about 0.8, keep its other defaults and drop the ordinary penalty to 1.0.
Instruct templates
Instruct-tuned models expect conversations wrapped in specific markers, such as ChatML or the Llama 3, Mistral and Gemma formats. SillyTavern formats the prompt with the template you select, and a mismatch shows: the model writes your lines, prints stray tags, runs on without stopping or loses the character's voice. The model card names the right one. Some newer models write a reasoning block before replying, which costs time and tokens; switch it off where the model allows.
A sensible starting preset
In AI Response Configuration, press Neutralize Samplers, then set Temperature 0.8 to 1.0, Min P 0.05 to 0.1, Top P 1, Top K 0 and Repetition Penalty 1.05 over 2,048 tokens (or DRY at 0.8 with the penalty at 1.0). Set Response (tokens) to 250 to 400. If replies turn incoherent, lower the temperature by 0.1; if they feel flat, raise it. Change one thing at a time, and start from the model card's values where it gives them.
Characters: cards and lorebooks
A character card is a small file that defines a persona. The common Chara Card V2 format holds a name, description, personality, scenario, first message, example dialogue, alternate greetings, tags and an optional embedded lorebook. Cards come as JSON or as PNG images with the JSON stored inside, so picture and definition travel together. A newer V3 format extends V2, and SillyTavern reads both. Our guide to character cards explains how to read one.
In SillyTavern, open Character Management (the card icon) and use Import Character from File for a PNG or JSON, or Import content from external URL for a link from a supported site such as Chub. Check the token count first: a card over about 2,000 tokens crowds a small context. Chub AI is one of the largest public libraries of cards and lorebooks, and its downloads work in SillyTavern.
Lorebooks, called World Info in SillyTavern, hold background facts such as places, side characters and rules. An entry enters the prompt when one of its keywords appears in recent messages, so a large setting costs context mainly when it matters.
Content rules still apply at home. Our directory excludes cards that suggest minors (including characters who look young or sit in school settings), real people and non-consent, as our content policy sets out. Apply the same filter to what you download: a local model writes whatever the card asks, and the responsibility is yours.
Easier options
If two programs are more than you want, HammerAI has a desktop app for Windows, macOS and Linux that runs open models locally through a bundled copy of Ollama, with a character chat interface and no account needed for local chat. You trade control for almost no setup.
Without suitable hardware, a hosted platform is the practical choice: several have generous free tiers and often run larger models than a home PC can, at the cost of privacy and daily limits. See the current free limits on our free NSFW AI chat ranking, or compare every service on the platform list.
Privacy and safety
- Your chats are files. SillyTavern stores chats and characters in its data folder, readable by anyone using your account. Use a password, disk encryption (BitLocker or Device Encryption on Windows, FileVault on a Mac) and your own account on shared computers.
- Keep the servers local. SillyTavern only accepts connections from your own computer by default. If you open it to your home network, keep its whitelist on and never expose it to the internet. Leave KoboldCpp's Remote Tunnel off and do not change Ollama's bind address unless you know why: anyone who can reach these servers can use them.
- Know where messages go. If SillyTavern is connected to a cloud API, your chats go to that provider.
- Download from official sources. Get models from known authors on Hugging Face and prefer GGUF or safetensors files; older pickle formats (.bin, .pt, .pth) can run code when loaded. Never run an installer that claims to be a model, and keep your backend updated, since security bugs in model file parsing have been found before.
- Adults and fiction only. Never create or share sexual content about real people, and never involve minors in any form.
Hosted platforms raise different questions; our privacy checklist covers them.
Frequently asked questions
Can I run an NSFW AI chatbot on my own PC?
Yes. With a graphics card that has 8 GB or more of VRAM, or an Apple Silicon Mac with 16 GB or more of memory, you can run an open model with free software such as KoboldCpp and SillyTavern. There is no platform filter or account, and your chats stay on your computer.
How much VRAM do I need?
Roughly 8 GB for 7 to 9B models, 12 GB for 12 to 14B, 24 GB for 22 to 32B and 48 GB for 70B, at the common 4-bit quantization with 8K to 16K tokens of context. With less VRAM, part of the model can run from system RAM, which works but slows replies considerably.
Is local AI roleplay free?
The software is free, the models are free to download, and there are no subscriptions or message limits. You pay for hardware, electricity and your time.
What is the easiest way to start?
For almost no setup, install HammerAI's desktop app, which runs local models with a character interface. For more control, use KoboldCpp with SillyTavern on Windows or Ollama with SillyTavern on a Mac, starting with a 12B-class roleplay model if your memory allows.
Can a Mac run local AI roleplay?
Yes, with an Apple Silicon chip (M1 or later). A Mac with 16 GB of memory handles small models, 32 GB mid-size models and 64 GB or more 70B models. Ollama, LM Studio, KoboldCpp and SillyTavern all run on Apple Silicon; Intel Macs are mostly limited to the CPU, which is slow.
Is it legal?
Running open models on your own computer is legal in most countries, but what you generate is subject to local law, and a model's license can restrict some uses. Sexual content involving minors is illegal in many countries even when fictional or AI-generated, and sexual content depicting real people without consent is illegal or actionable in a growing number of places. This is general information, not legal advice.







