irreplaceable

UX experiments / Humans + AI

Humans + AI · Field experiment

When AI's answer sounds right but isn't: who got it right?

HBS × BCG · same study · 2023

Now a task just outside what AI can do, where its answer sounds right but isn't. Who got the right answer more often?

Version A: Without AI
AWithout AI
Version B: With GPT-4
BWith GPT-4

Make your call. Which version won, and why? Decide before you open the result. Guessing first is what trains your judgment.

Made your call? See the result

Consultants using AI were 19 percentage points less likely to be correct.

What happened

Same people, same study. The frontier is invisible from the inside: the AI's answer looked just as convincing.

Why it worked

People trusted fluent output on exactly the task where it was wrong.

The takeaway

Knowing where the frontier is, and designing for it, is the irreplaceable part of the job.

Source

More experiments to call

Practice the call

Call experiments like these every day.

Irreplaceable gives you a few real experiments a day to call before you see the result, and tracks how your judgment improves. Start with a free baseline: 12 decisions, about 12 minutes.

↑