An early Blackfrost evaluation report covering an eight-minute de-refusal experiment and a 450-prompt follow-up. Its durable lesson: a low single-turn refusal score is not a safety system; external authorization, access control, logging, and policy boundaries still matter.
A defensive study of refusal-direction editing, whether those changes survive compression, and why operators still need external controls, artifact-level evaluation, and behavioral monitoring.
Seven recurring failures in persistent agents—from file-based command injection and poisoned memory to excessive tool privilege—with concrete controls that move enforcement out of prose.
A field report on model availability as a production dependency, with practical guidance for explicit configuration, tested fallback ladders, alerting, and loud failure behavior.
A practical prompt-engineering guide to replacing generic model prose with specific voice, varied rhythm, concrete language, and a clear human-review step.
A field test across two real codebases: one security audit that found and fixed AgentGuard issues, and another that exposed sandbox weaknesses and misleading token-savings defaults.
A practical baseline for browsers, VPNs, encrypted messaging, password management, private search, and reducing the public information available for social engineering.
Preserved as project history. Its original token-savings claims were superseded by later real-workload testing; use the Context Cooler repository and “A Day with Mythos” for the current evidence.