Product

Introducing PolicyLM-1.7B: a small, fast, open model that reads your content policy

Filip JankovicFilip Jankovic
October 6, 2026

PolicyLM-1.7B is for Trust & Safety teams that need to act on every message right away, such as in live chat, game lobbies, DMs, usernames, or anywhere a decision has to come back before the conversation moves on.

Fixed classifiers are fast and cheap, but they can't read your platform’s unique policy, so every rule change means annotating data and waiting for the model to be retrained. And if you’re using a commercial fixed classifier, you rely entirely on your vendor’s taxonomy and have no room to add your own labels or categories as you need to. On the flip side, an LLM will read your policy well, but most are too slow and expensive for live chat.

PolicyLM-1.7B is a decision model that gives you a good decision against your own policy and taxonomy right out of the box, at classifier speed and cost.

Because PolicyLM-1.7B reads your policy with every message, it's a good fit for teams whose rules don't map neatly onto anyone's off-the-shelf taxonomy, and where policy changes are frequent enough that you can’t wait on annotation and engineering for every update.

We’re releasing it as open weights (Apache 2.0), so it's free to download and use however you like. It's already proven itself in production: a custom fine-tuned version of PolicyLM runs on a platform that handles more than a million messages a day.

  • Fast. Under 100 ms, which is quick enough to keep a live conversation moving. (A median of 35 ms per short chat message on a 24 GB L4 with 6 categories.)
  • Cheap. Score all of your traffic instead of sampling it. Runs on a laptop or on a single 24 GB GPU.
  • Multi-label. Returns all labels and scores at once.
  • Accurate for its size. On our custom-policy benchmark, it beat every other model we ran under 20B parameters. PolicyLM-1.7B ships with two score cutoff presets for detection sensitivity. The default, "precision", suits live chat, where violations are rare. "Balanced" catches more, for queues where violations are common or when a miss costs more than a false flag. For recall-first triage, set a lower cutoff.
  • Policy-aware. You name the labels you want in plain language, as specific as you need, and your policy team can add or refine them with no model retraining needed. It can also label positive, pro-social content.
  • Text-only. Evaluated in 19 languages, with English strongest.
Chart of accuracy on custom policies against model size. PolicyLM-1.7B reaches about 83% accuracy; only gpt-oss-safeguard-20B and CoPE-B-A4B score as high or higher, and both are over 20B parameters.

The benchmark comparison is in the model card.

How it compares

Recently, TypeSafe AI launched Jev, and it's been getting a lot of attention in AI circles. Instead of writing out an answer the way a chatbot does, Jev just returns a decision, which makes it much faster and cheaper than a large LLM when a decision is all you need.

PolicyLM-1.7B is built on the same idea. If Jev caught your eye, PolicyLM-1.7B is the same kind of model, trained specifically for content moderation, that you can run yourself. It reads your policy and content together and returns a 0–1 score for each of your categories in a single pass, with no text generated. That's why it's fast and cheap, and why you get a score you can set a threshold on.

Its weights are also open, so you can try it, fine-tune it on your own community, and run it on your own infrastructure (or we’re happy to fine-tune and manage it for you, of course).

Big models are great for reasoning through appeals and nuanced policies, while smaller specialized models are great for analyzing and labeling content at scale. We expect teams to run several specialized models side by side or route from one to another, each doing what it's best at. That's why we built Musubi, and PolicyLM-1.7B is one piece of it.

Here's how PolicyLM-1.7B stacks up against the two options most teams use today, a fixed ML classifier and a larger LLM:

Fixed ML classifierLLMPolicyLM-1.7B
How it decidesScores content against categories set at training timeWrites its verdict one token at a time, based on a policyScores your policy in one pass; generates no text
What you get backA score for each of its built-in categoriesA text answer you parse, plus a reason if you ask for oneA 0–1 score for every category in your policy
SpeedTens of msUsually hundreds of ms or more<100 ms
Context windowVariesUsually long2048 tokens, policy and message together
Your policyNot read. Changing a rule means relabeling data and retraining, or waiting on the vendorRead from the prompt; can change any time, including what a label means. Cache rebuilding needed on policy change.Your categories and rules, read with every message. Policy edits take effect at once. No retraining and no cache to rebuild.
ThresholdsSet per categoryHard to tune on a text answerSet per category, with precision (default) and balanced presets
RecallStrong on what it was trained for; blind to anything outside its categoriesCan be prompted to catch more edge casesTuned for a balanced F1: a good answer you can act on at low latency, rather than maximum recall
ExplanationsA score, not a written rationaleCan write a rationaleA score, not a written rationale
Where it runsVendor API or your own infrastructureUsually a hosted APIOpen weights: your infrastructure or Musubi
Best forStable, well-defined harms that match its taxonomyAppeals, bans, takedowns, and novel judgment callsA decision on every message in live chat, DMs, and usernames, or when you can't wait

Limitations

  • Text only, one message at a time (no conversation history)
  • Tamil is the weakest of the 19 languages we evaluated, and every custom policy we tested was written in English
  • Benign content that sounds harmful, many unrelated categories in one call or long texts can raise false flags
  • No reasons provided

How PolicyLM-1.7B works in practice

On steerability

When people say a model reads your policy or is steerable, they can mean two different things:

  • Who picks the labels. PolicyLM-1.7B lets you define your own labels at any level of detail, compared to fixed classifiers which use one set taxonomy.
  • Who defines what a label means. PolicyLM-1.7B starts with meanings learned during training, then applies them to your labels and exceptions. These instructions can shift scores, but aren’t meant to redefine abuse as support. Some policy-adaptive models do allow labels to be redefined or inverted entirely, which can be useful flexibility, but makes it easier for user content or an accidental edit to undermine a rule. For enforcement, where the goal is consistent application of prohibitions at scale, an anchored model like PolicyLM-1.7B is usually safer.

So the short version is: PolicyLM-1.7B supports your taxonomy and your labels, with your definitions layered on top of ones the model already understands. You can write your policy as custom categories with short plain-language rules (or use the built-in Aegis taxonomy, NVIDIA's public list of 23 standard harm categories), send a message, and get back a 0–1 score for each category.

Download the repository and run the quickstart, in a fresh virtual environment: the helper needs transformers 4.57.6 (5.x is not supported yet).

Shell
pip install "huggingface_hub>=0.34,<1.0"
hf download musubilabs/policylm-1.7b --revision v1.2 --local-dir policylm-1.7b
pip install -r policylm-1.7b/inference/requirements.txt
python policylm-1.7b/inference/quickstart.py policylm-1.7b

Then score messages against your own policy:

Python
import sys; sys.path.insert(0, "policylm-1.7b/inference")
from policylm_infer import PolicyLM, Policy, Category

model = PolicyLM.from_pretrained("policylm-1.7b")
policy = Policy([
    Category("Harassment",
             violation_rule="Flag insults or threats aimed at another player.",
             not_violation_rule="Do not flag trash talk about the game itself or criticism of how someone played."),
    Category("Off-platform trading",
             violation_rule="Flag offers to trade or take payment outside the platform.",
             not_violation_rule="Do not flag warnings about scams or questions about the trading rules."),
])
r = model.classify("Selling 5k gold, pay me on PayPal first and I'll send it after.", policy)
print(r.violations, {c.name: round(c.score, 3) for c in r.categories.values()})

Set a threshold per category, starting from the “precision” (default) or “balanced” preset, and calibrate on a sample of your own content before you go live.

The helper cleans text before scoring: it undoes look-alike characters, spaced-out letters and base64, which caught 10 to 12 points more disguised violations on our test set (balanced cutoff) with no rise in false flags. Leetspeak mostly gets through. It also scores long messages in windows.

How we built it

PolicyLM-1.7B starts from BidirLM-1.7B-Embedding, an open encoder derived from Qwen3-1.7B-Base. We trained it on public safety datasets (NVIDIA's Nemotron Safety Guard Dataset v3, with some labels corrected with LLM help; Alibaba's XGuard-Train-Open-200K; PolyGuardMix) plus synthetic and LLM-written data.

Joining ROOST’s Model Community

We’re excited to announce that PolicyLM-1.7B is part of the ROOST Model Community!

To mark the launch, ROOST and Musubi are publishing Choosing and Routing Open Safety Models, an open source guide on how to build and select the right models for each task in safety. The guide answers this practical question: which model should handle which decisions when it comes to safety?

  • Part 1 covers selection through four questions: how accurate a model is on your policy, what it costs at your scale, how fast it is, and how easily you can steer it as your policy changes.
  • Part 2 covers routing: when to cascade from a small model to a larger one, how to set thresholds and escalation paths, and when a single model is the better call.

On October 27, ROOST and Musubi are hosting a workshop on safety models: “Does a Model Follow Your Rules?”. Bring a policy you work with and run it across RMC models, including PolicyLM, and we'll walk through where models agree, where they diverge, and how to set thresholds and routing for your own surfaces. Findings from the session will feed back into the RMC!

Get started

If you want to run PolicyLM-1.7B yourself, the weights and model card are on Hugging Face at musubilabs/policylm-1.7b. There’s also a hosted version on Baseten.

We can't wait to see what the T&S community does with it. Try it on your own content and tell us how it goes.

Wondering if Musubi could help?

We love to talk about the basics, the gnarly edge cases, and everything in between.

Let’s chat

Don’t miss a post

By subscribing you agree to our Privacy Policy