Delphi × Team Abraham · Conversation teardown

What actually happened inside the 63

A re-engagement sequence brought 63 dormant users back to Jay-I between 10–16 September. The coaching held up. Everything around it did not.

63 conversations · 157 threads · 1,830 messages in window · all hand-graded
Analysis 16 Sep 2026 · read-only snapshot delphi-full-2026-09-16.json

Team Abraham emailed four re-engagement messages to 1,117 cold Jay-I users. Sixty-three came back and used the product, generating 907 new messages — a 5.6% take-up, with the 13 September scarcity email pulling 31 of the 63 on its own.

That quantitative picture was already known. This is the qualitative one: every conversation read and graded on five dimensions, with every platform defect counted rather than anecdotally reported.

33%
left with a committed next action
21 of 63
32%
never got past the opening message
20 of 63
76
hard platform failures, hitting 24% of users
two distinct bugs
0.79
correlation between what the user typed and how good the answer was
Pearson r
01

Jay-I works. The campaign feeding it does not.

The single strongest predictor of answer quality is how much the user typed — not anything about Jay-I. Context richness and answer specificity move together almost in lockstep, on a clean monotonic ladder:

User's context richnessConversationsMean answer specificity
1 — vague generalities141.93
252.60
3103.10
4213.86
5 — numbers, named customers, real constraints134.85
The finding

Jay-I is not the bottleneck. Getting the user to talk is the bottleneck. Every intervention should be aimed there.

Where the coaching is genuinely good

Methodology fidelity averages 3.27 across all 63 — but 4.4 among the 33 conversations that got past the opener. Where users engaged, Jay-I correctly applied the three ways to grow a business, risk reversal, host-beneficiary and JV structures, and leverage of existing assets.

  • A gutter-cleaning business with 3,000 customers and no referral system — Jay-I located the exact moment of maximum goodwill (job finished, customer present) and built the ask there.
  • A medical-equipment firm stuck at $400k — Jay-I spotted that the risk reversal was backwards, with the client bearing the risk, and restructured it.
  • A retailer at 150 customers/month — when the user asked for more ideas, Jay-I refused and forced a single test instead.
Hold on. You came in saying you're not following up with customers at all, which is costing you real money. That's the thing to fix first, not to add new ideas on top of it.Jay-I, declining to chase a topic-hop

That refusal behaviour recurs throughout the corpus and is the most Jay-like thing in it.

02

The 76 failures are two bugs, not one

The platform errors reported to Delphi from anecdote are real and measurable. They are also not a single fault — they split cleanly into a deterministic input-size ceiling and a discrete service outage.

BugCountSignatureEvidence
A — input-size ceiling44 still liveLong paste → instant error100% failure above 4,000 chars (33/33)
B — outage, 10 Sep 19:08–19:31 UTC27 reported resolvedNormal input, 4 users at once82% failure rate in window vs 3.5% baseline
C — residual / unclassified5 6%1,000–2,000 char band
100% 75% 50% 25% 0% 3.5% 1.7% 15% 85% 100% 100% 100% <500 500–1k 1k–2k 2k–4k 4k–8k 8k–16k 16k+ n=771 n=58 n=26 n=13 n=18 n=9 n=6 Characters in the user's message
Bug A is deterministic. Median user input before an error: 3,179 characters. Before a normal reply: 134. Above 4,000 characters, 33 out of 33 attempts failed.

It is not a conversation-length problem. Errors on normal-length input occurred at a lower median cumulative thread size (3,884 chars) than successful turns (8,016) — which rules out context length and isolates Bug B as a genuine service incident.

This bug selects against your best users

The people who pasted 4,000–15,000 characters were the ones who had done the homework.

  • One pasted a 6,872-character consultancy brief. Killed instantly. Total session: four messages.
  • One pasted an 11,522-character prompt, errored twice, returned the next day, errored again, and gave up with "hi".
The most expensive single finding

One user pasted the campaign's own "CHAIN-BREAKER DIAGNOSTIC (Copy-Paste Ready)" prompt — 15,349 characters — four times. It errored every time. They never had a conversation at all.

The campaign shipped a copy-paste prompt that exceeds the platform's input ceiling.

And the error message itself compounds the damage. A 4,000-character paste returns:

Delphi is having trouble accessing my digital mind right now. Could you please repeat what you just said?58 occurrences

That is the worst possible response, because it invites the user to paste the same oversized text again — which fails again. Users did exactly that, up to four times in a row.

Two findings that clear Delphi

Both should be reported alongside the faults, because they change what the remaining complaints mean.

  • Duplicate greeting: 0 occurrences in 157 threads. This is an artefact of cloud actions, not UTM-coded entry; this cohort arrived via UTM links, so the bug could not present here. Not evidence it is fixed — only that it is scoped to the cloud-actions path.
  • Genuine content repetition: 0 occurrences. 93 near-duplicate agent turns were detected and every one was the error string repeating. The "Jay-I repeats itself" complaint is a symptom of Bug A and B, not a separate defect. Fix the errors and it disappears.
03

Why nine people opened it and left immediately

This was the highest-leverage question in the brief. The answer is not disengagement.

57 of 62 conversations (92%) opened with one of exactly three canned prompts. Only five people typed their own first message. The most-used opener appeared 32 times:

I want to find the untapped opportunities in what I'm already doing. I'll give you a messy rundown of my business, my audience, and where I'm stuck…Canned prompt · 32 of 62 conversations

That prompt writes a cheque the user has to cash. It promises a messy rundown the user has not written. Jay-I then correctly replies "I'm listening, give me the rundown" — and a cold user who clicked a button out of curiosity is facing a blank box and ten minutes of real work.

User's entire contributionJay-I's replyResult
canned prompt"I'm listening. Give me the messy version…"gone
canned prompt"Give me the rundown, and I'll ask what I need."gone
canned prompt"What's your current customer acquisition cost, and how does it compare to LTV…"gone
canned promptSpanish greeting to an English prompt, then "I'm listening. Walk me through it."gone
canned promptplatform error (10 Sep outage)gone

Three kill mechanisms, in order of size: effort asymmetry (the button promised something the user hadn't written); opening with a hard question a re-engaged user doesn't have to hand; and a platform error on the very first turn.

The instructive counter-example

One user pressed the same canned prompt — but had 40 prior interactions, so Jay-I had context and led with five named, specific assets of that user's business instead of a question. A genuinely impressive answer. That user still left after one message.

Value-first is necessary. For a re-engagement cohort it may not be sufficient on its own.

04

Scores across five dimensions

Activation
3.32
5 ones · 12 fives
Context
3.22
14 ones · 13 fives
Specificity
3.41
6 ones · 15 fives
Fidelity
3.27
6 ones · 10 fives
Outcome
2.62
21 ones · 6 fives

Outcome is the weakest dimension by a distance, and it is the one most clearly ours to fix. A third of all conversations trail off mid-air. Sessions end on an unanswered Jay-I question 44 times; 61 of 63 end on a Jay-I turn rather than a user's. Jay-I asks one more question instead of closing the loop.

Correlations — Specificity↔Fidelity r=0.93 · Activation↔Outcome r=0.85 · Context→Specificity r=0.79 · Context→Activation r=0.77 · Context→Outcome r=0.73. Specificity and Fidelity at 0.93 are close to measuring the same thing: when Jay-I gets concrete it is also at its most Jay-like. Worth collapsing in a future rubric.

05

What people actually asked about

User messages only, canned prompts excluded. This feeds the next campaign's copy and tells us which programme sections to load.

ClusterPeopleMentions
Scaling / capacity / systems26202
Lead generation / traffic19171
Referrals / JV / partnerships14167
Offer / pricing / packaging27131
Positioning / USP / messaging13106
Conversion / closing2178
Existing customer base / reactivation873
Starting from scratch1220
Personal / confidence / direction819

Referrals and JV punch far above their headcount — 14 people but 167 mentions, meaning the few who reach it go deep. That is core Jay territory and the current sequence never mentions it.

The gap is louder: reactivating an existing customer base — the most reliable Jay lever and the literal subject of the "What you already own, recombined" email — was raised by only eight people. The email did not convert its own theme into conversation.

06

All 63, graded

Anonymised. A Activation · C Context · S Specificity · F Fidelity · O Outcome, each 1–5.

Sort:
IDMsgsACS FOTotalErrWhat happened
07

What to do

Ask of Delphi — in priority order

  • Publish the input limit and fail gracefully. The current message invites the user to repeat the exact input that just failed. This alone accounts for 44 of 76 failures, and it is the one still live — do not let it be absorbed into the resolved upstream incident.
  • Confirm the resolved upstream incident covers 10 Sep 19:08–19:31 UTC, and ask for its scope — we can only see this cohort.
  • Note that "repeats itself" and "duplicate greeting" are, respectively, an artefact of the input ceiling and scoped to cloud actions — so the ceiling carries more reputational damage than it appears to.

Ours to fix — by expected return

  • Replace the canned opener. It demands effort before delivering value. Pre-fill a ten-second prompt, or open with something specific about the user.
  • Cap paste length client-side at ~1,500 characters with a visible counter, until the ceiling is fixed.
  • Retire the CHAIN-BREAKER copy-paste prompt or split it — at 15,349 characters it cannot succeed on the current platform.
  • Fix the close. A third of sessions end mid-air; Outcome is the weakest dimension in the dataset.
  • Pin the language to what the user typed, not their locale. 53 Italian, 35 German and 2 Japanese turns went to users who wrote in English.
  • Stop Jay-I claiming to have read URLs it has not read. The only outright hostile exit in the corpus — "Forget it, this was useless." — followed exactly that.
Campaign-level

The 13 September email threatened a deadline that was never enforced — 1,598 users remain on unlimited access. It pulled 31 of the 63. Re-using scarcity without enforcing it will cost credibility with precisely the segment that responded to it.