INSIGHTS

Statistics, explained for decisions in the real world.

Practical guidance on designing stronger studies, evaluating programs and AI systems, and avoiding the statistical traps that make confident conclusions fragile.

Visit the DASS blog

EXPLORE THE LIBRARY

More practical statistics, minus the jargon.

AI Evaluation

The 1979 IBM Warning That AI Still Hasn't Answered

A decades-old training-manual slide argued a computer can never be held accountable, so it must never make a management decision. Two 2025-2026 studies now show, in quantified detail, exactly why that gap hasn't closed.

AI Evaluation

The AI Raters That Scored Like Humans and Judged Nothing Like Them

Nine LLMs rating hate speech correlated with human judgments at 0.911 — a number that would pass almost any validation check. Decomposed into its parts, the two scales disagreed on severity, item calibration, question order, and who the comment was about.

Data Literacy

The Protective Effect That Wasn't There

Hospital records seemed to show diabetes protecting patients from gallbladder disease. The correlation was never real — it was manufactured by the simple fact that everyone in the data had already been admitted. A case study in Berkson's Paradox, and why a sample built by selection isn't safe to read at face value.

Research Practice

Building Publication-Ready Statistical Figures

A figure that misrepresents your own data's uncertainty undermines a paper more than an awkward color choice ever will. Here are the rules that actually affect whether a figure is trustworthy, not just tidy.

Research Practice

Writing Reproducible R Code for Academic Research

A model that only runs on your machine, in your session, with your undocumented fixes, isn't reproducible — and journals and funders are increasingly asking for proof that it is. Here's what actually makes R code reproducible.

Data Literacy

The Tallest Parents Never Had the Tallest Children

In 1886, Francis Galton compared thousands of parents' and children's heights and found the pattern he'd first spotted in sweet peas: every extreme pulled back toward average, generation after generation. A real case study in regression to the mean, and why picking the worst (or best) score to intervene on can make almost anything look like it worked.

Program Evaluation

When Does a Program Need a Control Group?

There's no participant-count threshold — a program needs a control group the moment you want to claim it caused an outcome, not just that the outcome happened. Here's how to tell, and what to do when a true control group isn't possible.

Data Literacy

The Correlation That Was Positive for States and Negative for People

In 1950, a sociologist found that immigrants and literacy correlated positively across U.S. states and negatively across individual people — the same census, the same year, opposite signs. A real case study in the ecological fallacy, and why a group-level number can't answer a question about the people inside it.

Study Design

When Does a Pilot Study Need to Become a Full RCT?

There's no headcount threshold. The moment anyone starts treating a pilot's numbers as evidence the intervention worked, the question has outgrown the design meant to answer it — here's how to tell, and what the next stage actually requires.

Data Literacy

The Survival Rates That Improved While No One Got Better

In 1985, lung cancer survival statistics rose in every single stage of the disease — without a single patient living longer. A real case study in the Will Rogers phenomenon, and why a rising average can hide the fact that nothing actually changed.

Study Design

When Is a Sample Too Small to Bother Analyzing?

There's no universal cutoff, but there are principled ways to tell before you run a single test — and a few situations where a small sample is still worth analyzing honestly.

Consulting

How Much Does a Statistical Consultant Actually Cost?

Rates range from a few hundred to several thousand dollars depending on scope, not on how complicated your data feels to you. Here's what actually drives the price, and how to get a quote you can trust.

Program Evaluation

What Funders Actually Expect From a Program Evaluation

Funders don't just want to know your program happened. They want evidence it caused the outcome you're claiming — and evaluations that skip that distinction tend to lose credibility exactly when it matters.

AI Evaluation

How to Evaluate an LLM Beyond Average Accuracy

A single accuracy number can hide the exact failures that matter most — subgroup gaps, instability across runs, and confident wrong answers. Here's what a more defensible evaluation actually measures.

Study Design

What Belongs in a Statistical Analysis Plan?

A statistical analysis plan written before you see the data is one of the cheapest ways to protect a study's credibility — and reviewers, funders, and IRBs increasingly expect to see one.

Study Design

The Antidepressant That Worked 94% of the Time — On Paper

Published trials made a class of antidepressants look almost uniformly effective. The FDA had the trials that never got published, and the real picture was a coin flip. A case study in publication bias, and why the literature you can read is not the literature that was run.

AI Evaluation

The Essay That Said Nothing and Scored a Perfect 6

An automated essay-grading engine gave top marks to student essays that made no sense at all, as long as they used long sentences and fancy words. A real case study in Goodhart's Law, and why optimizing a proxy metric isn't the same as measuring the thing you actually care about.

AI Evaluation

The Wolf Detector That Never Looked at the Wolf

A machine-learning model hit strong accuracy telling wolves from huskies — until researchers checked what it was actually looking at. A real case study in shortcut learning, and why a good accuracy number can hide exactly how a model is cheating.

Data Literacy

The Dead Salmon That Read Minds

A dead fish, put in a brain scanner, appeared to respond to photos of people with statistically significant brain activity. It was a joke with a serious point — about what happens when you run thousands of tests at once and only report the ones that worked.

Data Literacy

The Test Result That Fooled the Doctors

Give a room full of physicians a positive screening test and one prevalence number, and most will badly overestimate the odds the patient is actually sick. A short primer on base-rate neglect, and why 'accurate' tests aren't the same as trustworthy results.

Data Literacy

The Plane That Never Made It Back

The Air Force wanted to armor the parts of returning bombers riddled with bullet holes. A statistician told them to armor the parts that had none. A short primer on survivorship bias, and why the data you don't have can matter more than the data you do.

Data Literacy

The Bias That Disappeared When You Looked Closer

Berkeley's 1973 admissions numbers looked like textbook discrimination against women — until anyone checked department by department. A short primer on Simpson's Paradox, and why the number you aggregate is a choice, not a neutral fact.

Study DesignProgram Evaluation

The Program That Looked Like It Worked — Until Someone Checked

Scared Straight felt like it worked: kids were shaken, parents saw a difference, the numbers looked good. The best evidence says it made things worse. Here's why a good story keeps fooling careful people, and what to ask for instead.

HAVE A QUESTION OF YOUR OWN?

Turn the next statistical question into a defensible decision.

Start a conversation