
Expanding Our Uncensored Models - Five More That Decline on Capability, Not Policy
Five uncensored text models, charged per token against the balance you already have.
Today we are expanding the Uncensored section of PayPerQ with five uncensored text models, served directly by their upstream provider, Venice, rather than through an intermediary. Each one is built on open weights whose trained refusal behaviour has been reduced or removed, so they answer what they are able to answer rather than what a content policy allows. There is no subscription and no separate account: they are charged per token against the balance you already have.
Most AI models have two different reasons for not answering you, and they look identical from the outside. Sometimes the model genuinely cannot do the thing. Sometimes it can do the thing perfectly well and has been trained not to. You get the same apologetic paragraph either way, usually somewhere in the middle of the work you were doing. These five models are for the second case.
Why an uncensored section at all
Mainstream assistants refuse far more than the law or common sense requires. A fiction writer gets cut off mid-scene. A role play breaks character to deliver a content warning nobody asked for. A security researcher asking how an attack works gets a lecture about why attacks are bad. None of that is the model running out of ability — it is behaviour that was trained into it, and what a provider's guidelines say is allowed and what its model will actually do are not always the same thing. You only ever meet the second.
We have offered one uncensored model for a while. One row is a gesture, not an option: it gives you nothing to choose between and no price range. Five rows is a choice.
We are also not deciding for ourselves which models qualify. The upstream catalogue carries an uncensored flag per model, and we use it as the qualification rule: anything in our Uncensored section has to carry that flag. It is a floor, not a shopping list — more models upstream carry the flag than we have picked up. A section header is a promise, and we would rather keep it narrow than have it be approximately true.
The five models
| Model | Context |
|---|---|
| Venice Uncensored 1.2 | 128K |
| Venice Role Play Uncensored | 128K |
| Gemma 4 Uncensored | 256K |
| GLM 4.7 Flash Heretic | 200K |
| Gemma 4 26B A4B Uncensored | 64K |
Charged per token against your balance — no subscription and no minimum. Current rates are on the pricing page and in the model picker.
Venice Uncensored 1.2 is the flagship and the one to start with, built with the Dolphin team on Mistral's 24B foundation. Venice Role Play Uncensored is tuned specifically for staying in character over long conversations. GLM 4.7 Flash Heretic is an uncensored variant of GLM 4.7 Flash, the cheapest of the set, and supports discounted cached reads, which makes it the sensible default for long sessions where you keep resending the same context. Gemma 4 Uncensored and Gemma 4 26B A4B Uncensored are uncensored derivatives of Google's Gemma 4 — the base model is served elsewhere in our catalog; these derivatives are not.
You will find them in the Uncensored section of the model picker, below the private TEE models.
Images in the conversation
Venice Uncensored 1.2 and Gemma 4 Uncensored can also read images you attach: up to 10 per message on Venice Uncensored 1.2, and one per message on Gemma 4 Uncensored. Gemma 4 Uncensored only sees the most recent image in the conversation, so attach again whatever you want it to compare. PNG, JPEG and WebP are supported. The other three models are text-only.
Web search
All five can search the web, run upstream rather than routed through a separate search provider. The control sits in the composer and is Off or On, with no automatic middle setting — you decide, per conversation — and it ships Off. Turns that search carry a small surcharge on top of tokens; a turn that searches nothing costs nothing extra.
How a model loses its refusals
A mainstream model is trained twice. First it learns language and knowledge from an enormous amount of text. Then it goes through safety training that teaches it to decline whole categories of requests. That second step is why jailbreak prompts are fragile: the refusal is not a setting sitting in front of the model, it is in the weights.
There are two common ways to take it back out:
- Fine-tuning. The model is trained again on data from which refusals and moralising have been filtered out, until declining stops being its reflex. This is the approach behind the flagship of this set.
- Abliteration. Researchers locate the internal direction the model uses when it is about to refuse and edit it out of the weights directly, without a full retraining run.
What neither method changes is what the model knows. Knowledge, reasoning and context length are those of the base model underneath. An uncensored model is still a model: it hallucinates, it has a knowledge cutoff, and on hard reasoning it will lose to the frontier models we already list. Use it for the work where the refusal was the problem, not as a general upgrade.
An uncensored model is not an uncensored platform
The weights are only one layer. The service in front of them is the other, and it can undo everything the weights allow — a restrictive system prompt injected into every conversation, or a filter on the output, and an uncensored model starts refusing again.
So here is what we do on these five. We add no content filter of our own and no system prompt of our own. The upstream provider would otherwise layer its own system prompt on top of every request; we switch that off explicitly, so the model behaves the way its weights do rather than the way an upstream prompt we did not write would steer it.
Uncensored is not the same as lawless. The provider's terms still apply, our Terms of Service still apply, and content that is illegal to create is illegal whichever model wrote it. What changes is who decides what is worth asking — you, not a refusal trained in by someone else.
One note on the enclave model
Gemma 4 26B A4B Uncensored runs inside a hardware enclave upstream. Like the other four models it is zero data retention — prompts and responses are not stored after the request completes — but that is a retention promise, not an encryption one. Our TEE models are end-to-end encrypted in your browser and we can prove it; this model is routed over the ordinary chat path, and we make no attestation claim for it. If encrypted inference is what you need, use the TEE models — that is what they are for.
Three things uncensored does not mean
It does not mean smarter. Removing refusals removes refusals. The model underneath is exactly as capable, and exactly as fallible, as it was before.
It does not mean private. Censorship and privacy are separate questions, and a platform can get one right and the other wrong. All five of these models are zero data retention upstream, but that is a property of how they are served, not of being uncensored. If you want to understand the difference between the privacy levels we offer, read Anon vs ZDR vs TEE.
It does not mean a jailbroken mainstream model. A jailbreak argues with weights that were trained to refuse. The output tends to be hedged and inconsistent, and the trick stops working the next time the provider patches it. These models have nothing to argue with.
FAQ
Is it legal to use an uncensored model? For legitimate work — fiction, research, security testing, personal projects — generally yes. What you generate is still subject to the law and to our Terms. The model changes what it is willing to write, not what is permitted.
Why does one of these models still decline sometimes? Because the trained refusal was reduced, not every limit the model has. It may simply not know the answer, and some behaviour from the base model can survive the process. Rephrasing usually helps; so does trying another of the five.
Which one should I start with? Venice Uncensored 1.2 for general use. Venice Role Play Uncensored for characters and long-running stories. GLM 4.7 Flash Heretic when cost matters and the context is long. Gemma 4 Uncensored when you need the 256K context window.
Can I use them through the API?
Yes. They are ordinary models on the PayPerQ API, addressed by id — for example venice/venice-uncensored-1-2 — with the same API key and the same balance.
A real choice
Five more models does not make uncensored inference the right choice for every prompt. It makes it a real choice, with a price range and an actual decision behind it, instead of a single row you either used or ignored.
They are available now. Open the Uncensored section in the model picker, or address any of these models through the PayPerQ API.
What is PayPerQ?

PayPerQ is a pay-per-query AI service that gives you instant access to hundreds of chat, image, video, and audio AI models in one place. Unlike traditional ChatGPT subscription that charges $20+ per month, PPQ users pay only for what they use—averaging just $4 a month.
Account registration optional. No monthly commitments. Privacy focused. Credit cards and all major cryptos accepted. Just top up with as little as 10 cents and start using premium AI immediately.
Why Use PayPerQ?
- Access hundreds of AI models from all major providers in one place
- Pay per use - no subscriptions, no wasted money on unused credits
- No registration required - start using AI in seconds
- Start small - top up with as little as 10 cents
- Privacy-first - conversational data stored locally by default
- Average cost: ~1 cent per query
Getting Started
- Visit ppq.ai
- Top up your balance (crypto or credit card, as little as 10 cents)
- Open the model selector, scroll to the Uncensored section, and pick a model
No account creation needed—just fund and go.