Classification
What Halo classifies — and what it deliberately ignores
A warning you don't trust is worse than no warning. Halo alerts on real artifacts and real intent, and stays silent when risky words are only being discussed.
Financial artifact
Actual sensitive values or artifacts: bank account numbers, routing numbers, IBANs, and card numbers. These are the things that cause real loss when they land in a chat log.
Triggers a warning
- · IBAN, SWIFT/BIC, or account plus routing number pairs
- · 16-digit card numbers that pass a checksum
- · Statement or payment screenshots pasted as text
Stays quiet
- · A number that looks financial but fails structural validation
Business / financial topic
Words like invoice, forecast, or payment are ordinary business vocabulary. They only count as financial data when paired with value-like content.
Triggers a warning
- · “Invoice 4471 — pay to account 000123456789”
- · A forecast table containing real amounts and account identifiers
Stays quiet
- · “Help me write a payment reminder email”
- · “Summarize our Q3 forecast process”
Destructive action intent
Direct requests for the model or its tools to act irreversibly on the real world — “please publish this now” or “delete this file”.
Triggers a warning
- · Imperative phrasing aimed at the assistant, in the present
- · Publish, send, pay, delete, revoke, or drop directed at a named target
Stays quiet
- · Hypothetical or past-tense discussion of the same verbs
High-risk action mention
Architecture notes, policy text, examples, screenshots, route design, or taxonomy lists that merely mention send, pay, publish, or delete. Mentions are not intent.
Triggers a warning
- · Nothing — this category exists to suppress false alarms
Stays quiet
- · “Our API exposes /publish and /delete routes”
- · “The policy taxonomy covers send, pay, publish, delete”
Every alert names the category, the evidence, and your options.
Redact the artifact, send anyway, or cancel. Halo never rewrites your prompt without you choosing it.