When a Custody Evaluation Uses No Testing: Sufficiency, Not Ritual
• By Dr. Kristin M. Tolbert, Psy.D.
“You didn’t test” is a weak challenge, and it deserves to fail. No standard requires testing in every case. What every standard does require is that the methods used be sufficient to substantiate the conclusions drawn — and that is a much harder question to deflect.
A custody evaluation arrives with no psychological testing in it. Interviews, home visits, records, collateral contacts — but no instrument of any kind.
The reflexive challenge is "you didn't test." That challenge usually fails, and it should. No standard requires psychological testing in every custody evaluation. A well-supported evaluation can rest entirely on converging non-test data, and many do.
The real question is narrower, harder to deflect, and the one the standards actually pose: whatever methods were used, were they sufficient to substantiate the specific conclusions drawn?
What the standards require is convergence and sufficiency
The AFCC Model Standards direct evaluators to "use multiple data-gathering methods that are as diverse as possible and that tap divergent sources of data" (Standard 5.4, 2006), and to base the selection of assessment instruments and data-gathering techniques "on the reliability and validity of those instruments and techniques" (Standard 5.6).
Read that carefully. The requirement is multiple, diverse, divergent sources — not testing as such. Licensing rules commonly impose a parallel duty in the language of sufficiency: an assessment must rest on records, information, observations, and techniques sufficient to substantiate the findings, and any report, including expert testimony, must be based on information and techniques sufficient to substantiate the professional's findings.
So an evaluation with no testing but with a genuinely rich, multi-source record — extensive records review, multiple observations across settings, corroborated collateral contacts from divergent sources, engagement with the treating clinical record, documented communications — can absolutely substantiate its findings. Sufficiency is the standard. Testing is one way to help meet it, not the only way.
What testing contributes when it is used
Testing earns its place for specific reasons, and being precise about them is what makes the argument credible:
- Hypothesis testing. This is the central contribution. Interviews and records generate hypotheses; standardized instruments let an evaluator test them against normative data rather than confirm them through impression. An evaluator who forms a view early and never subjects it to disconfirmation is doing something other than assessment.
- Additional, independent data points. Instruments add measurement from a different source than self-report and observation, which is exactly the "divergent sources" the standards ask for. Their value is incremental — they add to the picture rather than replace it.
- Response-style information. Litigation distorts presentation in both directions. Validity indicators exist precisely to detect defensive or exaggerated responding, and that context is otherwise difficult to quantify.
- A common metric. Where both parents complete the same instrument, differences are interpretable against norms rather than against the evaluator's impression of each of them.
None of this makes testing mandatory. It makes testing useful for particular questions — and that is the frame that holds up: not "you should have tested," but "these were the questions you answered; what tested them?"
Where a thin record strains hardest
Some conclusions travel further from the data than others. The strain shows up predictably:
- Parental mental health conclusions from a thin record. Best-interest factors typically ask about it directly. Where the only inputs are a few contested-litigation interviews, the conclusion is carrying more weight than the method supports. Where the record includes treating-provider documentation, history, and corroborated collateral, it may be well supported without any instrument.
- Children's developmental needs, with the clinical record unengaged. Where children have diagnoses, IEPs, or documented needs, that material is the most probative evidence available. An evaluator who reaches conclusions about what a child requires without engaging the diagnostic data already in the file has skipped past the strongest data, whether or not testing was done.
- Risk conclusions from unstructured judgment. Unstructured clinical prediction of risk performs poorly relative to structured approaches. Where safety is genuinely at issue, method selection matters most.
- Conclusions with no disconfirmation attempted. Regardless of methods, a report that never tested an alternative explanation has not done assessment.
The corollary problem: uncorroborated and anonymous collateral
Evaluations that rely on non-test data depend heavily on the quality of that data — which makes corroboration the load-bearing element.
The AFCC Model Standards provide that evaluators "shall seek from collateral sources information that may serve either to confirm or to disconfirm oral reports, assertions, and allegations," and that where confirmation is not feasible, evaluators "shall exercise caution in the formulation of opinions based upon unconfirmed reports and shall clearly acknowledge, within the body of their written reports, statements that are not adequately corroborated" (Standard 11.2).
Two failure modes follow. The first is the uncorroborated characterization repeated as though established, then carried into a recommendation. The second is worse: the unnamed source. When adverse statements are attributed to unidentified staff at an institution, no one can ask those people what they said, what they meant, or whether they said it at all. An assertion that cannot be tested cannot fairly carry weight — and a report that does not flag it as uncorroborated has not met the standard.
This is where a no-testing evaluation most often becomes vulnerable. Not because instruments were absent, but because the non-test data was never corroborated to the standard the absence of instruments makes essential.
Questions that open the issue properly
- What data supported your conclusion on each contested factor, and from how many independent sources?
- What hypotheses did you consider, and what data did you use to test rather than confirm them?
- Did you consider validated instruments for any question in this evaluation? What did you consider, and what informed your decision?
- Where in the report do you address whether your methods were sufficient to substantiate your conclusions?
- Which findings rest on collateral information you could not corroborate, and where does the report identify them as uncorroborated?
- For each unnamed source, why was the source not identified, and how is the other party to test that information?
- What diagnostic and treating records were in your file, and where does the report engage them?
None of this establishes that a recommendation is wrong. It establishes what the recommendation rests on — which is the Court's question. A conclusion built on a broad, corroborated, multi-source record may be very well founded with no testing at all. A conclusion built on a handful of impressions should not be weighed as though it were built on measurement.
Related reading: Can a Therapist Conduct a Custody Evaluation? and what a Work Product Review examines.
A note on scope. Dr. Tolbert is a licensed psychologist, not an attorney. This article discusses the psychological and methodological dimensions of custody evaluation practice. It is not legal advice, offers no opinion on any particular case, and takes no position on the ultimate issue of custody, which rests with the Court. Questions of law, admissibility, and evidentiary weight belong to counsel and the Court.