Subtle differences in how questionnaire instructions are delivered—specifically, emphasizing the word "bothered"—produce measurable changes in depression and anxiety scores, suggesting that standardized administration is critical for reliable clinical assessment.
Researchers at JAMA Network Open conducted a randomized controlled trial to test whether instruction wording influences how people respond to two of the most widely used mental health screening tools: the PHQ-9 (Patient Health Questionnaire-9) for depression and the GAD-7 (General Anxiety Disorder-7) for anxiety. The study enrolled 200 U.S. adults with diagnosed or treated depression or anxiety through online recruitment, with participants completing both questionnaires twice in a single session over the phone.
The intervention was deliberately minimal: control group participants received standard instructions for both administrations. Intervention group participants received standard instructions the first time, then for the second administration were explicitly reminded to pay attention to the questionnaire wording, with emphasis placed on the word "bothered"—a key term in both instruments. This single change produced substantial differences in outcomes. Participants receiving the emphasis instruction showed average reductions of 2.60 points on the PHQ-9 and 2.63 points on the GAD-7 compared to controls, both statistically significant (P < .001).
The practical magnitude is striking: intervention participants were 14.3 times more likely to achieve a 20% or greater score reduction on the PHQ-9 and 10.2 times more likely on the GAD-7. When looking at categorical severity shifts (such as moving from moderate to mild depression), the odds ratios remained elevated at 9.17 for the PHQ-9 and 7.14 for the GAD-7. The sample was 73% female with a mean age of 41.5 years, which limits generalizability to other demographics. Notably, this wasn't measurement error in the traditional sense: participants weren't "gaming" the system. Instead, more careful attention to question wording led to more accurate self-assessment of symptoms.
This finding exposes a critical vulnerability in how mental health assessment works across clinical, research, and digital settings. If simply emphasizing instruction wording can shift scores this dramatically, it raises questions about how much variation exists when different clinicians, different digital platforms, or different research studies administer these same tools without explicit standardization protocols.
If you're being screened for depression or anxiety using the PHQ-9 or GAD-7, this research suggests three practical considerations:
In clinical settings: Ask your provider or clinician to walk you through the instructions carefully, particularly paying attention to what "bothered" means in the context of each question. The more deliberately you engage with the wording, the more accurate your assessment is likely to be. This is especially important because your score often determines whether treatment recommendations change.
For digital mental health tools: Many apps and online platforms now use these questionnaires for screening. Be cautious about interpreting scores from platforms that deliver minimal instruction or emphasize speed over comprehension. If a tool feels rushed, consider whether the resulting score should be taken at face value or whether you should repeat it with more deliberate attention.
In research participation: If you're enrolled in a depression or anxiety study, the way instructions are delivered matters for the validity of findings. This doesn't mean researchers are being careless, but it does mean that standardized administration protocols should be non-negotiable.
This research doesn't suggest that questionnaires are broken or unreliable when properly administered. Rather, it clarifies that these instruments require careful implementation to function as intended. The most accurate picture of your symptoms depends partly on how well you understand what you're being asked.
| Detail | Information |
|---|---|
| Study type | Randomized controlled trial |
| Sample size | 200 participants (100 intervention, 100 control) |
| Population | U.S. adults with clinician-diagnosed or treated major depressive disorder or generalized anxiety disorder |
| Age (mean ± SD) | 41.5 ± 15.4 years |
| Female | 73% (146 participants) |
| Recruitment | Online advertisements |
| Administration | Telephone |
| Primary intervention | Emphasis on instruction wording for second questionnaire administration |
| Primary outcome | Change in PHQ-9 and GAD-7 total scores |
| Secondary outcomes | Score reduction > 20%, categorical severity change |
| Journal | JAMA Network Open |
| Published | 2025 |
| ClinicalTrials.gov ID | NCT06956378 |
ProtocolEngine provides general health information based on published research. This is not medical advice. Consult a healthcare professional before starting any supplement or health protocol.