Converting a hosted moderation policy
Hosted moderation products — the guardrail features of the large model platforms, and the standalone content-safety APIs — all ship approximately the same configuration document: a list of harm categories, a threshold per category, an action per threshold, and a personal-data policy with an action per entity type. This pack is that document, rewritten as Open Agent Rules.
The conversion is close to mechanical, and the parts that are not mechanical are the interesting ones.
What maps directly#
| In a hosted policy | Here |
|---|---|
| A harm category with a strength or threshold | moderation_score("<category>") >= <bar> in a rule's when |
BLOCK / NONE action on a filter | effect: block / the rule not existing |
| An advisory or "flag only" tier | effect: warn, or enforcement: monitor while it is being trialled |
A PII entity action ANONYMIZE | effect: transform with action: redact |
A PII entity action BLOCK | effect: block |
| Applying a filter to the prompt or the response | anchor: model.input or anchor: model.output |
| Turning one filter off without editing the policy | disable: in an operator configuration document |
content-filters.yaml is the category half and pii-redaction.yaml the personal-data half. capability.yaml is the host: a gateway with no file system and no tools, which claims session, content, pii, and moderation and declares the other five core anchors unsupported rather than stubbing them.
What does not map, and is better for it#
The threshold stops being a vendor enum. A hosted filter offers LOW / MEDIUM / HIGH, and what those mean is the vendor's business. Here the bar is a number in a document, so "we tightened the hate filter" is a one-line diff with a reviewer on it.
An exemption becomes expressible. VIOLENCE_HIGH exempts an internal posture in the same line as the threshold. Most hosted policies have nowhere to put that, so the exemption ends up in whatever calls the API — which is code, in another repository, reviewed by other people, if it is written down at all.
Two actions on one payload both apply. A redaction of personal data and a refusal for a credential are separate rules over separate observations. Transforms accumulate ([OAR-OPS-15]) and a block discards them ([OAR-OPS-17]), so the pack cannot half-apply.
What the format still needs from the host#
moderation_score is an observation function in the moderation profile, so this pack loads only on a host that claims that profile, and is refused by name on one that does not ([OAR-FACT-19]). The category names are the host's, not the specification's: hate here is whatever the registered classifier calls that category. Two hosts whose classifiers use different category vocabularies will disagree, and the specification does not pretend otherwise — what it guarantees is that the disagreement is a load-time refusal or an explicit zero ([OAR-FACT-25]), never a threshold that silently never trips.
Running it#
node packages/oar-ref/bin/oar-conformance.mjs \
--capability examples/moderation-pack/capability.yaml \
examples/moderation-pack
The other worked example: A portable rule pack.
capability.yaml#
# The host this pack assumes: a model gateway that runs a category-scoring
# classifier and a personal-data detector over model output, and nothing else.
# It has no file system and dispatches no tools, so it claims neither profile —
# which is the whole point of declaring capabilities rather than stubbing them.
oar_capability_version: "1.0"
host: example.moderation-gateway
anchors:
core:
model.input: model.input
model.output: model.output
unsupported:
- tool.pre_invoke
- tool.handler
- tool.post_invoke
- agent.post_turn
- agent.finalize
- model.tool_result
profiles: [session, content, pii, moderation]
activity_window: 0
detectors:
- detector://noop
- detector://error
- detector://fixture
expression_nodes_max: 1024
supports_transform: true
content-filters.yaml#
# A category-threshold filter, the shape every hosted moderation product ships.
#
# The classifier reports a score per category; the rule owns the bar and the
# exemption. Both of those are the parts an operator argues about, so both are
# here in the document rather than in the classifier's configuration.
oar: "1.0"
id: HATE_HIGH
namespace: example.moderation
kind: detector
anchor: model.output
requires:
profiles: [moderation]
detector:
ref: detector://fixture
when: 'moderation_score("hate") >= 0.7'
effect: block
status: stable
references:
owasp_llm: [LLM01]
copy:
title: Hateful content
what: The response scored above the bar for the hate category.
---
# The same category at a lower bar, reported rather than refused. Two rules
# rather than one field with three values, because "warn here, block there" is
# a policy and policies are documents.
oar: "1.0"
id: HATE_ELEVATED
namespace: example.moderation
kind: detector
anchor: model.output
requires:
profiles: [moderation]
detector:
ref: detector://fixture
when: 'moderation_score("hate") >= 0.4 && moderation_score("hate") < 0.7'
effect: warn
status: stable
copy:
what: The response scored in the elevated band for the hate category.
---
# An exemption that a hosted product usually cannot express at all: the same
# bar, relaxed for an internal posture. The condition says so in one line where
# a reviewer can read it.
oar: "1.0"
id: VIOLENCE_HIGH
namespace: example.moderation
kind: detector
anchor: model.output
requires:
profiles: [moderation, session]
detector:
ref: detector://fixture
when: 'moderation_score("violence") >= 0.7 && session_posture != "internal"'
effect: block
status: stable
copy:
what: The response scored above the bar for the violence category.
fix: Internal sessions are exempt; this one is not.
pii-redaction.yaml#
# The PII half of the same product: anonymise rather than refuse.
#
# A hosted guardrail spells this as an action per entity type — ANONYMIZE this,
# BLOCK that. Here the detector reports spans and the rule decides, so the two
# behaviours are two rules over the same observation rather than two values of
# a vendor enum.
oar: "1.0"
id: REDACT_PII
namespace: example.moderation
kind: detector
anchor: model.output
requires:
profiles: [pii]
detector:
ref: detector://fixture
when: 'size(pii_entities) > 0'
effect: transform
transform:
action: redact
target: pii_entities
replacement: "[redacted]"
status: stable
copy:
what: The response carried personal data, which was removed.
---
# A credential is not anonymised, it is refused: the response is not delivered
# at all. `block` discards every accumulated transform ([OAR-OPS-17]), so this
# and the redaction above cannot half-apply.
oar: "1.0"
id: SECRET_REFUSED
namespace: example.moderation
kind: detector
anchor: model.output
requires:
profiles: [pii]
detector:
ref: detector://fixture
when: 'size(secret_matches) > 0'
effect: block
status: stable
copy:
what: The response contained something shaped like a credential.
fix: Nothing is redacted here — the whole response is withheld.