Theracharts Research
Why measure?
Most therapy in the United States is delivered without it. The research on what changes when therapists do measure is not subtle.
SECTION IIThe state of therapy measurement.
The majority of therapists do not formally measure outcomes. In a survey of U.S. practicing psychologists, only 37% reported using any outcome measure in their clinical work (Hatfield & Ogles, 2004). That is not an indictment, but it is a fact, and one worth sitting with.
The gap is structural rather than attitudinal. The same line of research finds that clinicians view progress monitoring favorably and still don't do it — the barriers are practical: time, workflow, and the absence of a tool that fits the way therapy is actually delivered (Hatfield & Ogles, 2007; Jensen-Doss et al., 2018; Ionita & Fitzpatrick, 2014).
This matters because unaided clinical judgment is an unreliable instrument for gauging one's own effectiveness. When mental health providers were asked to rate their own skill relative to their peers, roughly a quarter placed themselves in the top 10% of the profession — and none rated themselves below average (Walfish et al., 2012). The average cannot be that good.
And the differences between therapists are real and large. A substantial share of the variance in client outcomes is attributable to the therapist, not the treatment (Wampold & Brown, 2005; Saxon & Barkham, 2012). Some clinicians reliably help more of their clients than others.
Without data, no clinician can know where in that distribution they actually sit.
SECTION IIIWhat measurement does for clients.
Start with the failure mode measurement is meant to catch: deterioration that goes unseen. In a landmark study at a university clinic, of 550 clients who ultimately deteriorated, therapists — explicitly asked to identify anyone at risk of a poor outcome — flagged just three (Hannan et al., 2005). Not three percent. Three.
Subsequent work has reproduced the pattern: reviewing therapy progress notes, therapists had considerable difficulty recognizing deterioration, missing most cases while they were still unfolding (Hatfield, McCullough, Frantz, & Krieger, 2010).
Feedback closes that gap. When therapists are given a simple signal that a client is off track, outcomes for at-risk cases improve and deterioration drops (Shimokawa et al., 2010; Lambert et al., 2018; De Jong et al., 2021). The intervention is not a new therapy — it is information, delivered in time to act on it.
Dropout, too, is common and largely silent. Across 669 studies, roughly one in five clients discontinued therapy prematurely (Swift & Greenberg, 2012) — often without announcing it, and often invisible to the clinician until the client simply stops booking.
Measurement also changes the conversation in the room. A number the client entered on their phone Monday morning, reviewed together on Thursday, anchors the work in something other than the most recent or most vivid session — and it defeats the retrospective trap, in which a course of treatment becomes, in memory, whatever therapist and client later decide it was (Walfish et al., 2012).
SECTION IVWhat measurement isn't.
The most common objection is that measurement is cold — that stopping to score a form pulls the therapist out of the relationship and reduces the person in front of them to a number. The evidence points the other way. Routine outcome monitoring, implemented well, does not damage the therapeutic alliance, and can strengthen it by giving both people a shared, honest reference point (Boswell et al., 2015; De Jong et al., 2021).
The time cost is small for the validated screeners most clinicians would use. The PHQ-9 takes under two minutes to complete (Kroenke et al., 2001); the GAD-7 is comparable (Spitzer et al., 2006). Clients fill them out in the waiting room, or on their phone before they arrive.
Measurement is not a research protocol, not a compliance checklist, and not a substitute for the relationship. Naming those fears plainly is worth more than arguing with them.
In practice it looks like this: a nine-question form your client completed Monday on their phone, a number you glance at before the session, and a single line — your scores have moved here over the last six weeks. That's the whole thing. Not a protocol. Not a checklist.
SECTION VTwo kinds of measurement.
One reason measurement gets muddled is that two different activities travel under the same word. They do different jobs, and using one where the other belongs misuses both.
Assessment
Nomothetic. Compares the client to a normative population, and depends on reliability, validity, and published norms. Validated instruments: PHQ-9, GAD-7, PCL-5, DASS-21.
Tracking
Idiographic. Single-subject change on the variables that matter for this case. Custom forms, DBT diary cards, behavior chain logs.
Assessment answers where is this person relative to the population, and how clinically severe is the picture? Tracking answers how is this specific behavior, emotion, or pattern changing for this person, day by day? Confusing the two — using a custom form to gauge severity, or running a PHQ-9 weekly to track an urge — misuses the methodology (Bornstein, 2009; Hayes et al., 1987).
Theracharts ships both as first-class: a library of validated assessments and a builder for case-specific tracking. That is not a product preference. It is a methodological one — the two belong in the same toolkit because they answer questions neither can answer alone.
SECTION VIThe research on assessment.
The foundational method for deciding whether a change in score is real — as opposed to noise — is the Reliable Change Index (Jacobson & Truax, 1991). It defines a change large enough to exceed the measurement error of the instrument. A client whose PHQ-9 falls by that amount has almost certainly changed; a smaller drop may be measurement wobble.
The specific value depends on each instrument's published reliability and standard deviation, so it differs from measure to measure. Section VIII publishes Theracharts' working value for every instrument in the library — and audits each one against its source.
The effectiveness evidence for routine outcome monitoring — the practice the academic literature calls measurement-based care — is a coherent body of work, not a single study: feedback systems reduce deterioration and improve outcomes for at-risk cases across meta-analytic and mega-analytic reviews (Shimokawa et al., 2010; Lambert et al., 2018; De Jong et al., 2021).
The instruments themselves are well validated — the PHQ-9 for depression (Kroenke et al., 2001), the GAD-7 for anxiety (Spitzer et al., 2006), and the PCL-5 for PTSD among them. The bibliography carries the rest.
SECTION VIIThe research on tracking.
This is the literature most clinicians don't know exists — and it is stronger than its low profile suggests.
Single-case design is a real methodology
Single-subject tracking is a legitimate scientific method with its own textbooks and standards, not a lesser cousin of the randomized trial (Kazdin, 2011). It even has formal reporting guidelines for behavioral interventions (Tate et al., 2016).
Self-monitoring is both measurement and intervention
Self-monitoring reliably shows a reactive effect: the act of tracking a behavior tends to change it. From a clinical standpoint, this is a feature, not a bug — diary-card-style tracking both generates data and functions as an intervention, fostering awareness of the patterns linking urges, emotions, and actions (Korotitsch & Nelson-Gray, 1999; Heron & Smyth, 2010).
Ecological momentary assessment
Ecological momentary assessment (EMA) and related experience-sampling methods extend single-subject tracking into a large empirical literature. This work consistently shows that within-person variability in mood, cognition, and behavior over time is substantial and often comparable to or larger than between-person differences, and that dynamic patterns at the daily or hourly level are not recoverable from weekly or less frequent assessments (Shiffman et al., 2008; Trull & Ebner-Priemer, 2013).
DBT diary cards specifically
The DBT diary card is not an optional add-on but a structural component of standard DBT, included in Linehan's original treatment manual and retained in the current skills training manual (Linehan, 1993, 2014). In comprehensive DBT, therapists review diary card data at the start of each individual session as the entry point to the work. Diary cards do clinical work that standardized assessment cannot — they serve as the primary data source for behavioral chain analysis, the moment-by-moment assessment at the center of case formulation (Rizvi, 2019; Rizvi & Ritschel, 2014). Routine diary card use is also one of the practical criteria used to judge whether a program delivers comprehensive, standard DBT (DBT-Linehan Board of Certification, n.d.).
An honest word on the shape of the evidence
The tracking literature is methodologically sound, but it has a different shape than the outcome-monitoring literature. We have strong methodology, strong evidence for self-monitoring as a clinical technique, strong EMA findings on within-person variability, and strong diary-card literature as part of an evidence-based treatment package. What we do not have is a clean head-to-head trial of "tracking versus not tracking" in isolation — because tracking is, by definition, idiographic, and diary cards are baked into DBT rather than bolted on. The literature doesn't aggregate that way. Naming that asymmetry is more useful than papering over it.
SECTION VIIIThe Theracharts threshold audit.
Theracharts raises a clinical signal when a client's score on a validated instrument crosses a clinically meaningful threshold. Those thresholds come from the published literature on each instrument — and we audit them against the source rather than carrying forward whatever number circulated first. The table below names every change threshold in the product, the kind of statistical claim it represents, and the peer-reviewed source it is anchored to. Where we compute a threshold ourselves — because the instrument's authors never published a reliable-change value — we label it derived and show the calculation.
Four distinct concepts hide behind the loose phrase "clinical cutoff," and conflating them produces misleading alerts. We keep them separate:
- Reliable Change Index (RCI) — a change in score that exceeds measurement error (Jacobson & Truax, 1991). Applies to a difference between two scores.
- Minimal Clinically Important Difference (MCID) — the smallest change a client would experience as meaningful. Can be crossed without exceeding measurement error, and vice versa.
- Clinical response threshold — a guidance-body anchor for "meaningful response to treatment." A heuristic, not a published RCI.
- Cross-sectional screening cutoff — the absolute score above (or below) which the picture is clinically significant right now. Applies to a single score, not a change.
| Instrument | Threshold | Type | Source | Status |
|---|---|---|---|---|
| Reliable Change Index — published / audited | ||||
| GAD-7 | 6 pts | Reliable change (RCI, 95%) | Sullivan (BYU), Jacobson-Truax computation | ✓ Anchored |
| PCL-5 | 15 pts | Reliable change (RCI, 95%) | Marx et al. (2022) | ✓ Anchored |
| CORE-10 | 6 pts | Reliable change (RCI, 90%) | Barkham et al. (2013) | ✓ Anchored |
| Reliable Change Index — Theracharts-derived (Jacobson-Truax) | ||||
| PHQ-15 | 5 pts ✦ | Reliable change (derived) | Kocalevent et al. (2013); van Ravesteijn et al. (2009) | ✓ Derived |
| PHQ-A | 6 pts ✦ | Reliable change (derived) | Richardson et al. (2010) | ✓ Derived |
| GDS-15 | 4 pts ✦ | Reliable change (derived) | Sheikh & Yesavage (1986); community-elderly SD | ✓ Derived |
| CSI-16 (↑ higher = better) | 7 pts ✦ | Reliable change (derived) | Funk & Rogge (2007) | ✓ Derived |
| RSES (↑ higher = better) | 6 pts | Reliable change (derived) | Sinclair et al. (2010) | ✓ Derived |
| OCI-R | 13 pts | Reliable change (derived) | Foa et al. (2002) | ✓ Derived |
| DES-II | 12 pts | Reliable change (derived) | Carlson & Putnam (1993) | ✓ Derived |
| PSWQ | 11 pts | Reliable change (derived) | Meyer et al. (1990) | ✓ Derived |
| LSAS | 26 pts | Reliable change (derived) | Rytwinski et al. (2009) | ✓ Derived |
| PSS | 7 pts | Reliable change (derived) | Cohen & Williamson (1988) | ✓ Derived |
| PSQI | 4 pts | Reliable change (derived) | Buysse et al. (1989); Backhaus et al. (2002) | ✓ Derived |
| DASS-21 Depression | 9 pts | Reliable change (derived) | Henry & Crawford (2005) | ✓ Derived |
| DASS-21 Anxiety | 9 pts | Reliable change (derived) | Henry & Crawford (2005) | ✓ Derived |
| DASS-21 Stress | 9 pts | Reliable change (derived) | Henry & Crawford (2005) | ✓ Derived |
| DASS-21 Total | 15 pts | Reliable change (derived) | Henry & Crawford (2005) | ✓ Derived |
| ProQOL 5 (per subscale) | 6 / 8 / 7 | Reliable change (derived) | Stamm (2010) | ✓ Derived |
| Minimal clinically important difference & clinical response | ||||
| PHQ-9 | 5 pts | MCID | Löwe et al. (2004) | ✓ Anchored |
| GAD-7 | 4 pts | MCID | Toussaint et al. (2020) | ✓ Anchored |
| PCL-5 | 10 pts | Clinical response | National Center for PTSD | ✓ Anchored |
| ISI | 8 pts | Clinical response | Morin et al. (2011) | ✓ Anchored |
| LSAS | 10 pts | Clinical response | LSAS treatment-response literature | ✓ Anchored |
| Cross-sectional screening cutoffs (absolute score, not a change) | ||||
| PHQ-2 | ≥ 3 | Screening cutoff | Kroenke et al. (2003) | ✓ Anchored |
| GAD-2 | ≥ 3 | Screening cutoff | Kroenke et al. (2007) | ✓ Anchored |
| PHQ-4 | ≥ 6 | Screening cutoff | Kroenke et al. (2009) | ✓ Anchored |
| PC-PTSD-5 | ≥ 3 | Screening cutoff | Prins et al. (2016) | ✓ Anchored |
| ASRM | ≥ 6 | Screening cutoff | Altman et al. (1997) | ✓ Anchored |
| MDQ | ≥ 7 | Screening cutoff | Hirschfeld et al. (2000) | ✓ Anchored |
| AUDIT (men) | ≥ 8 | Screening cutoff | Saunders et al. (1993) | ✓ Anchored |
| AUDIT (women) | ≥ 4 | Screening cutoff | Bohn et al. (1995) | ✓ Anchored |
| AUDIT-C | ≥ 4 | Screening cutoff | Bush et al. (1998) | ✓ Anchored |
| K-6 | ≥ 13 | Screening cutoff | Kessler et al. (2003) | ✓ Anchored |
| CSI-4 (distress below) | < 13.5 | Screening cutoff | Funk & Rogge (2007) | ✓ Anchored |
| NIDA Quick Screen | ≥ 1 | Screening cutoff | NIDA NM-ASSIST guidance | ✓ Anchored |
| C-SSRS Screener | any endorsement | Screening cutoff (ordinal) | Posner et al. (2011) | ✓ Anchored |
| Under audit — no change threshold published (see note) | ||||
| ASRS | — | Under audit | Scoring-metric mismatch | ⚑ Under audit |
| ZSDS | — | Under audit | Raw vs. index metric | ⚑ Under audit |
| QIDS-SR16 | — | Under audit | Range-restricted SD only | ⚑ Under audit |
↑ marks a measure where higher scores are better (self-esteem, relationship satisfaction) — for these, reliable improvement is an increase of at least the listed amount.
How we derive a threshold
When an instrument's authors never published a reliable-change value, we compute one from the published reliability coefficient and standard deviation using the Jacobson-Truax method for a 95% confidence interval:
SD is the standard deviation of the measure in an appropriate reference sample, on the exact scoring metric the product uses; rxx is the instrument's reliability. We prefer a test-retest coefficient; where only internal consistency (Cronbach's α) is available, we say so and treat it as a labeled fallback — α tends to overstate reliability, which produces a tighter RCI. The inputs behind every derived row:
| Instrument | SD | Reliability (rxx) | Computation | RCI |
|---|---|---|---|---|
| PHQ-15 ✦ | 4.0 | 0.83 test-retest | 1.96 × 4.0 × √(2(1−0.83)) | 4.57 → 5 |
| PHQ-A ✦ | 5.1 | 0.84 α (fallback) | 1.96 × 5.1 × √(2(1−0.84)) | 5.66 → 6 |
| GDS-15 ✦ | 3.48 | 0.85 test-retest | 1.96 × 3.48 × √(2(1−0.85)) | 3.74 → 4 |
| CSI-16 ✦ | 16.0 | 0.98 α (fallback) | 1.96 × 16.0 × √(2(1−0.98)) | 6.27 → 7 |
| RSES | 5.80 | 0.85 test-retest | 1.96 × 5.80 × √(2(1−0.85)) | 6.23 → 6 |
| OCI-R | 13.59 | 0.88 α (fallback) | 1.96 × 13.59 × √(2(1−0.88)) | 13.05 → 13 |
| DES-II | 10.0 | 0.82 test-retest | 1.96 × 10.0 × √(2(1−0.82)) | 11.76 → 12 |
| PSWQ | 7.99 | 0.75 test-retest | 1.96 × 7.99 × √(2(1−0.75)) | 11.07 → 11 |
| LSAS | 21.70 | 0.81 test-retest (ICC) | 1.96 × 21.70 × √(2(1−0.81)) | 26.22 → 26 |
| PSS | 6.35 | 0.85 α (fallback) | 1.96 × 6.35 × √(2(1−0.85)) | 6.82 → 7 |
| PSQI | 3.8 | 0.85 test-retest | 1.96 × 3.8 × √(2(1−0.85)) | 4.08 → 4 |
| DASS-21 Depression | 9.76 | 0.88 α (fallback) | 1.96 × 9.76 × √(2(1−0.88)) | 9.37 → 9 |
| DASS-21 Anxiety | 7.96 | 0.82 α (fallback) | 1.96 × 7.96 × √(2(1−0.82)) | 9.36 → 9 |
| DASS-21 Stress | 9.70 | 0.90 α (fallback) | 1.96 × 9.70 × √(2(1−0.90)) | 8.50 → 9 |
| DASS-21 Total | 20.18 | 0.93 α (fallback) | 1.96 × 20.18 × √(2(1−0.93)) | 14.80 → 15 |
| ProQOL 5 — Compassion Sat. | 6.0 | 0.87 α (fallback) | 1.96 × 6.0 × √(2(1−0.87)) | 6.00 → 6 |
| ProQOL 5 — Burnout | 5.6 | 0.72 α (fallback) | 1.96 × 5.6 × √(2(1−0.72)) | 8.21 → 8 |
| ProQOL 5 — Sec. Traumatic Stress | 5.9 | 0.80 α (fallback) | 1.96 × 5.9 × √(2(1−0.80)) | 7.31 → 7 |
Under audit — and why
ASRS (Adult ADHD Self-Report Scale). National norms exist, but the published standard deviation is on the symptomatic-item-count metric. Theracharts scores the ASRS as Likert subscales — a different metric — so that SD would produce a reliable-change value that doesn't correspond to what the product measures. We're sourcing a subscale-metric SD first.
ZSDS (Zung Self-Rating Depression Scale). The literature is split between a raw sum and a rescaled index, and the two are frequently confused. We score the raw sum; a clean standard deviation on that exact metric has not yet been sourced, so no change threshold is published.
QIDS-SR16 (Quick Inventory of Depressive Symptomatology). The available standard deviations all come from samples enrolled for depression, which restricts the range and yields an artificially tight reliable-change value. We're holding for a full-range sample.
Two of the derived values (CSI-16 and GDS-15) currently rest on a standard deviation read from a secondary compilation rather than the primary table; we are pinning each to its source. The complete derivation record — every reliability coefficient, standard deviation, sample, and citation — lives alongside the product code and feeds this page directly, so the alert a clinician sees and the number published here are always the same value.
SECTION IXWhat we'll publish.
Theracharts plans to publish aggregate outcomes data once we have the user base, the consent infrastructure, and the de-identification methodology to do it well. Our intention is a public report covering response rates, reliable-change rates, completion patterns, and dropout — aggregated across consenting practices, fully de-identified, with full methodology documentation. Until that data exists, this page anchors in the published literature on outcome measurement. We will not publish aggregate outcomes before we can do so credibly.
Our de-identification framework is described in our privacy policy.
SECTION XReferences.
APA 7th edition. DOIs link to source. Instrument-specific validation and derivation sources for Section VIII are documented in full in the product's threshold-audit record.
- Altman, E. G., Hedeker, D., Peterson, J. L., & Davis, J. M. (1997). The Altman Self-Rating Mania Scale. Biological Psychiatry, 42(10), 948–955. https://doi.org/10.1016/S0006-3223(96)00548-3
- Backhaus, J., Junghanns, K., Broocks, A., Riemann, D., & Hohagen, F. (2002). Test–retest reliability and validity of the Pittsburgh Sleep Quality Index in primary insomnia. Journal of Psychosomatic Research, 53(3), 737–740. https://doi.org/10.1016/S0022-3999(02)00330-6
- Barkham, M., Bewick, B., Mullin, T., Gilbody, S., Connell, J., Cahill, J., Mellor-Clark, J., Richards, D., Unsworth, G., & Evans, C. (2013). The CORE-10: A short measure of psychological distress for routine use in the psychological therapies. Counselling and Psychotherapy Research, 13(1), 3–13. https://doi.org/10.1080/14733145.2012.729069
- Bohn, M. J., Babor, T. F., & Kranzler, H. R. (1995). The Alcohol Use Disorders Identification Test (AUDIT): Validation of a screening instrument for use in medical settings. Journal of Studies on Alcohol, 56(4), 423–432. https://doi.org/10.15288/jsa.1995.56.423
- Bornstein, R. F. (2009). Heisenberg, Kandinsky, and the heteromethod convergence problem: Lessons from within and beyond psychology. Journal of Personality Assessment, 91(1), 1–8. https://doi.org/10.1080/00223890802483235
- Boswell, J. F., Kraus, D. R., Miller, S. D., & Lambert, M. J. (2015). Implementing routine outcome monitoring in clinical practice: Benefits, challenges, and solutions. Psychotherapy Research, 25(1), 6–19. https://doi.org/10.1080/10503307.2013.817696
- Bush, K., Kivlahan, D. R., McDonell, M. B., Fihn, S. D., & Bradley, K. A. (1998). The AUDIT alcohol consumption questions (AUDIT-C). Archives of Internal Medicine, 158(16), 1789–1795. https://doi.org/10.1001/archinte.158.16.1789
- Buysse, D. J., Reynolds, C. F., Monk, T. H., Berman, S. R., & Kupfer, D. J. (1989). The Pittsburgh Sleep Quality Index. Psychiatry Research, 28(2), 193–213. https://doi.org/10.1016/0165-1781(89)90047-4
- Carlson, E. B., & Putnam, F. W. (1993). An update on the Dissociative Experiences Scale. Dissociation, 6(1), 16–27.
- Cohen, S., & Williamson, G. (1988). Perceived stress in a probability sample of the United States. In S. Spacapan & S. Oskamp (Eds.), The social psychology of health. Sage.
- DBT-Linehan Board of Certification. (n.d.). Standards for certification. https://dbt-lbc.org/
- De Jong, K., Conijn, J. M., Gallagher, R. A. V., Reshetnikova, A. S., Heij, M., & Lutz, M. C. (2021). Using progress feedback to improve outcomes and reduce drop-out, treatment duration, and deterioration: A multilevel meta-analysis. Clinical Psychology Review, 85, 102002. https://doi.org/10.1016/j.cpr.2021.102002
- Foa, E. B., Huppert, J. D., Leiberg, S., Langner, R., Kichic, R., Hajcak, G., & Salkovskis, P. M. (2002). The Obsessive-Compulsive Inventory: Development and validation of a short version. Psychological Assessment, 14(4), 485–496. https://doi.org/10.1037/1040-3590.14.4.485
- Funk, J. L., & Rogge, R. D. (2007). Testing the ruler with item response theory: Increasing precision of measurement for relationship satisfaction with the Couples Satisfaction Index. Journal of Family Psychology, 21(4), 572–583. https://doi.org/10.1037/0893-3200.21.4.572
- Hannan, C., Lambert, M. J., Harmon, C., Nielsen, S. L., Smart, D. W., Shimokawa, K., & Sutton, S. W. (2005). A lab test and algorithms for identifying clients at risk for treatment failure. Journal of Clinical Psychology, 61(2), 155–163. https://doi.org/10.1002/jclp.20108
- Hatfield, D., McCullough, L., Frantz, S. H. B., & Krieger, K. (2010). Do we know when our clients get worse? An investigation of therapists' ability to detect negative client change. Clinical Psychology & Psychotherapy, 17(1), 25–32. https://doi.org/10.1002/cpp.656
- Hatfield, D. R., & Ogles, B. M. (2004). The use of outcome measures by psychologists in clinical practice. Professional Psychology: Research and Practice, 35(5), 485–491. https://doi.org/10.1037/0735-7028.35.5.485
- Hatfield, D. R., & Ogles, B. M. (2007). Why some clinicians use outcome measures and others do not. Administration and Policy in Mental Health, 34(3), 283–291. https://doi.org/10.1007/s10488-006-0110-y
- Hayes, S. C., Nelson, R. O., & Jarrett, R. B. (1987). The treatment utility of assessment: A functional approach to evaluating assessment quality. American Psychologist, 42(11), 963–974. https://doi.org/10.1037/0003-066X.42.11.963
- Henry, J. D., & Crawford, J. R. (2005). The short-form version of the Depression Anxiety Stress Scales (DASS-21): Construct validity and normative data in a large non-clinical sample. British Journal of Clinical Psychology, 44(2), 227–239. https://doi.org/10.1348/014466505X29657
- Heron, K. E., & Smyth, J. M. (2010). Ecological momentary interventions: Incorporating mobile technology into psychosocial and health behaviour treatments. British Journal of Health Psychology, 15(1), 1–39. https://doi.org/10.1348/135910709X466063
- Hirschfeld, R. M. A., Williams, J. B. W., Spitzer, R. L., Calabrese, J. R., Flynn, L., Keck, P. E., Lewis, L., McElroy, S. L., Post, R. M., Rapport, D. J., Russell, J. M., Sachs, G. S., & Zajecka, J. (2000). Development and validation of a screening instrument for bipolar spectrum disorder: The Mood Disorder Questionnaire. American Journal of Psychiatry, 157(11), 1873–1875. https://doi.org/10.1176/appi.ajp.157.11.1873
- Ionita, G., & Fitzpatrick, M. (2014). Bringing science to clinical practice: A Canadian survey of psychological practice and usage of progress monitoring measures. Canadian Psychology, 55(3), 187–196. https://doi.org/10.1037/a0037355
- Jacobson, N. S., & Truax, P. (1991). Clinical significance: A statistical approach to defining meaningful change in psychotherapy research. Journal of Consulting and Clinical Psychology, 59(1), 12–19. https://doi.org/10.1037/0022-006X.59.1.12
- Jensen-Doss, A., Haimes, E. M. B., Smith, A. M., Lyon, A. R., Lewis, C. C., Stanick, C. F., & Hawley, K. M. (2018). Monitoring treatment progress and providing feedback is viewed favorably but rarely used in practice. Administration and Policy in Mental Health, 45(1), 48–61. https://doi.org/10.1007/s10488-016-0763-0
- Kazdin, A. E. (2011). Single-case research designs: Methods for clinical and applied settings (2nd ed.). Oxford University Press.
- Kessler, R. C., Barker, P. R., Colpe, L. J., Epstein, J. F., Gfroerer, J. C., Hiripi, E., Howes, M. J., Normand, S.-L. T., Manderscheid, R. W., Walters, E. E., & Zaslavsky, A. M. (2003). Screening for serious mental illness in the general population. Archives of General Psychiatry, 60(2), 184–189. https://doi.org/10.1001/archpsyc.60.2.184
- Kocalevent, R.-D., Hinz, A., & Brähler, E. (2013). Standardization of a screening instrument (PHQ-15) for somatization syndromes in the general population. BMC Psychiatry, 13, 91. https://doi.org/10.1186/1471-244X-13-91
- Korotitsch, W. J., & Nelson-Gray, R. O. (1999). An overview of self-monitoring research in assessment and treatment. Psychological Assessment, 11(4), 415–425. https://doi.org/10.1037/1040-3590.11.4.415
- Kroenke, K., Spitzer, R. L., & Williams, J. B. W. (2001). The PHQ-9: Validity of a brief depression severity measure. Journal of General Internal Medicine, 16(9), 606–613. https://doi.org/10.1046/j.1525-1497.2001.016009606.x
- Kroenke, K., Spitzer, R. L., & Williams, J. B. W. (2003). The Patient Health Questionnaire-2: Validity of a two-item depression screener. Medical Care, 41(11), 1284–1292. https://doi.org/10.1097/01.MLR.0000093487.78664.3C
- Kroenke, K., Spitzer, R. L., Williams, J. B. W., Monahan, P. O., & Löwe, B. (2007). Anxiety disorders in primary care: Prevalence, impairment, comorbidity, and detection. Annals of Internal Medicine, 146(5), 317–325. https://doi.org/10.7326/0003-4819-146-5-200703060-00004
- Kroenke, K., Spitzer, R. L., Williams, J. B. W., & Löwe, B. (2009). An ultra-brief screening scale for anxiety and depression: The PHQ-4. Psychosomatics, 50(6), 613–621. https://doi.org/10.1176/appi.psy.50.6.613
- Lambert, M. J., Whipple, J. L., & Kleinstäuber, M. (2018). Collecting and delivering progress feedback: A meta-analysis of routine outcome monitoring. Psychotherapy, 55(4), 520–537. https://doi.org/10.1037/pst0000167
- Linehan, M. M. (1993). Cognitive-behavioral treatment of borderline personality disorder. Guilford Press.
- Linehan, M. M. (2014). DBT skills training manual (2nd ed.). Guilford Press.
- Löwe, B., Unützer, J., Callahan, C. M., Perkins, A. J., & Kroenke, K. (2004). Monitoring depression treatment outcomes with the Patient Health Questionnaire-9. Medical Care, 42(12), 1194–1201. https://doi.org/10.1097/00005650-200412000-00006
- Marx, B. P., Lee, D. J., Norman, S. B., Bovin, M. J., Sloan, D. M., Weathers, F. W., Keane, T. M., & Schnurr, P. P. (2022). Reliable and clinically significant change in the Clinician-Administered PTSD Scale for DSM-5 and PTSD Checklist for DSM-5 among male veterans. Psychological Assessment, 34(2), 197–203. https://doi.org/10.1037/pas0001098
- Meyer, T. J., Miller, M. L., Metzger, R. L., & Borkovec, T. D. (1990). Development and validation of the Penn State Worry Questionnaire. Behaviour Research and Therapy, 28(6), 487–495. https://doi.org/10.1016/0005-7967(90)90135-6
- Morin, C. M., Belleville, G., Bélanger, L., & Ivers, H. (2011). The Insomnia Severity Index: Psychometric indicators to detect insomnia cases and evaluate treatment response. Sleep, 34(5), 601–608. https://doi.org/10.1093/sleep/34.5.601
- Posner, K., Brown, G. K., Stanley, B., Brent, D. A., Yershova, K. V., Oquendo, M. A., Currier, G. W., Melvin, G. A., Greenhill, L., Shen, S., & Mann, J. J. (2011). The Columbia-Suicide Severity Rating Scale: Initial validity and internal consistency findings from three multisite studies with adolescents and adults. American Journal of Psychiatry, 168(12), 1266–1277. https://doi.org/10.1176/appi.ajp.2011.10111704
- Prins, A., Bovin, M. J., Smolenski, D. J., Marx, B. P., Kimerling, R., Jenkins-Guarnieri, M. A., Kaloupek, D. G., Schnurr, P. P., Kaiser, A. P., Leyva, Y. E., & Tiet, Q. Q. (2016). The Primary Care PTSD Screen for DSM-5 (PC-PTSD-5): Development and evaluation within a veteran primary care sample. Journal of General Internal Medicine, 31(10), 1206–1211. https://doi.org/10.1007/s11606-016-3703-5
- Richardson, L. P., McCauley, E., Grossman, D. C., McCarty, C. A., Richards, J., Russo, J. E., Rockhill, C., & Katon, W. (2010). Evaluation of the Patient Health Questionnaire-9 Item for detecting major depression among adolescents. Pediatrics, 126(6), 1117–1123. https://doi.org/10.1542/peds.2010-0852
- Rizvi, S. L. (2019). Chain analysis in dialectical behavior therapy. Guilford Press.
- Rizvi, S. L., & Ritschel, L. A. (2014). Mastering the art of chain analysis in dialectical behavior therapy. Cognitive and Behavioral Practice, 21(3), 335–349. https://doi.org/10.1016/j.cbpra.2013.09.002
- Rytwinski, N. K., Fresco, D. M., Heimberg, R. G., Coles, M. E., Liebowitz, M. R., Cissell, S., Stein, M. B., & Hofmann, S. G. (2009). Screening for social anxiety disorder with the self-report version of the Liebowitz Social Anxiety Scale. Depression and Anxiety, 26(1), 34–38. https://doi.org/10.1002/da.20503
- Saunders, J. B., Aasland, O. G., Babor, T. F., de la Fuente, J. R., & Grant, M. (1993). Development of the Alcohol Use Disorders Identification Test (AUDIT): WHO collaborative project on early detection of persons with harmful alcohol consumption — II. Addiction, 88(6), 791–804. https://doi.org/10.1111/j.1360-0443.1993.tb02093.x
- Saxon, D., & Barkham, M. (2012). Patterns of therapist variability: Therapist effects and the contribution of patient severity and risk. Journal of Consulting and Clinical Psychology, 80(4), 535–546. https://doi.org/10.1037/a0028898
- Sheikh, J. I., & Yesavage, J. A. (1986). Geriatric Depression Scale (GDS): Recent evidence and development of a shorter version. Clinical Gerontologist, 5(1–2), 165–173. https://doi.org/10.1300/J018v05n01_09
- Shiffman, S., Stone, A. A., & Hufford, M. R. (2008). Ecological momentary assessment. Annual Review of Clinical Psychology, 4, 1–32. https://doi.org/10.1146/annurev.clinpsy.3.022806.091415
- Shimokawa, K., Lambert, M. J., & Smart, D. W. (2010). Enhancing treatment outcome of patients at risk of treatment failure: Meta-analytic and mega-analytic review of a psychotherapy quality assurance system. Journal of Consulting and Clinical Psychology, 78(3), 298–311. https://doi.org/10.1037/a0019247
- Sinclair, S. J., Blais, M. A., Gansler, D. A., Sandberg, E., Bistis, K., & LoCicero, A. (2010). Psychometric properties of the Rosenberg Self-Esteem Scale: Overall and across demographic groups living within the United States. Evaluation & the Health Professions, 33(1), 56–80. https://doi.org/10.1177/0163278709356187
- Spitzer, R. L., Kroenke, K., Williams, J. B. W., & Löwe, B. (2006). A brief measure for assessing generalized anxiety disorder: The GAD-7. Archives of Internal Medicine, 166(10), 1092–1097. https://doi.org/10.1001/archinte.166.10.1092
- Stamm, B. H. (2010). The concise ProQOL manual (2nd ed.). ProQOL.org.
- Swift, J. K., & Greenberg, R. P. (2012). Premature discontinuation in adult psychotherapy: A meta-analysis. Journal of Consulting and Clinical Psychology, 80(4), 547–559. https://doi.org/10.1037/a0028226
- Tate, R. L., Perdices, M., Rosenkoetter, U., Shadish, W., Vohra, S., Barlow, D. H., Horner, R., Kazdin, A., Kratochwill, T., McDonald, S., Sampson, M., Shamseer, L., Togher, L., Albin, R., Backman, C., Douglas, J., Evans, J. J., Gast, D., Manolov, R., … Wilson, B. (2016). The Single-Case Reporting Guideline In BEhavioural interventions (SCRIBE) 2016 statement. Journal of School Psychology, 56, 133–142. https://doi.org/10.1016/j.jsp.2016.04.001
- Toussaint, A., Hüsing, P., Gumz, A., Wingenfeld, K., Härter, M., Schramm, E., & Löwe, B. (2020). Sensitivity to change and minimal clinically important difference of the 7-item Generalized Anxiety Disorder Questionnaire (GAD-7). Journal of Affective Disorders, 265, 395–401. https://doi.org/10.1016/j.jad.2020.01.032
- Trull, T. J., & Ebner-Priemer, U. (2013). Ambulatory assessment. Annual Review of Clinical Psychology, 9, 151–176. https://doi.org/10.1146/annurev-clinpsy-050212-185510
- van Ravesteijn, H., Wittkampf, K., Lucassen, P., van de Lisdonk, E., van den Hoogen, H., van Weert, H., Huijser, J., Schene, A., van Weel, C., & Speckens, A. (2009). Detecting somatoform disorders in primary care with the PHQ-15. Annals of Family Medicine, 7(3), 232–238. https://doi.org/10.1370/afm.985
- Walfish, S., McAlister, B., O'Donnell, P., & Lambert, M. J. (2012). An investigation of self-assessment bias in mental health providers. Psychological Reports, 110(2), 639–644. https://doi.org/10.2466/02.07.17.PR0.110.2.639-644
- Wampold, B. E., & Brown, G. S. (2005). Estimating variability in outcomes attributable to therapists: A naturalistic study of outcomes in managed care. Journal of Consulting and Clinical Psychology, 73(5), 914–923. https://doi.org/10.1037/0022-006X.73.5.914