Llama Guard 3 on Cloudflare
Model: @cf/meta/llama-guard-3-8b. Worker: angels-meta-llama-guard-v1.
The endpoint POST /classify accepts a JSON object with text (up to 1,200 characters) and an optional mode of prompt or response. Inference is allowed only after Cloudflare Access authentication. The result is fail-closed: approved is true only when the model returns an unambiguous safe classification.
The atomic daily budget reserves no more than four model calls per UTC day through the Worker. This does not guarantee account-wide zero cost, since other AI workloads share the Cloudflare allowance. Monitor account-wide usage.
Acceptance must test authenticated safe and unsafe inputs, unauthenticated access denial, quota exhaustion, and unexpected model output.