For one week in March, an agent read every email that hit my inbox, drafted replies, and queued them for my approval. Rules of engagement: it could draft anything, send nothing. The gate stayed shut — I reviewed every word before it left.
The box score
- 61 replies drafted. I sent 44 untouched, edited 15, rewrote 2. That edit rate is better than some humans I've managed, and the humans didn't work at 6 a.m.
- 9 correctly flagged as "you need to actually think about this one." Knowing what not to answer is the skill. It had it.
- 1 near-incident: it drafted a genuinely gracious reply to an obvious invoice scam, thanking them for their patience. Polite to a fault. The fault was mine — my instructions said "be warm to vendors" and it obeyed. Prompt fixed, lesson logged: agents don't have judgment, they have your judgment, compiled.
What actually made it work
Not the model. The scaffolding. It had my tone examples, my "who is this person" context, and explicit escalation rules. In other words, the same onboarding I'd give a human assistant, minus the desk.
Net time saved: about four hours across the week, most of it in decision fatigue I didn't feel. The inbox stopped being a to-do list written by strangers and became a queue of pre-chewed choices. That's the honest pitch for agentic email in 2026 — not "never read email again," but "read it like an executive instead of a clerk."
Verdict: hired, with supervision, like everyone else on payroll. It even survived the only performance review that matters — the task got done.