Tree testing: a practical guide
What tree testing measures, how to design tasks and a tree, which metrics to read, and how to run a tree test in Versive.
Tree testing measures whether people can navigate using labels and structure alone, without visual design as a cue. Use it to evaluate a new site map, an app's settings menu, or a help center's category structure. The method requires realistic tasks, a representative tree, and analysis of both success and the path each participant took.
What tree testing measures
A tree test strips a navigation structure down to text: a list of nested categories and items, with no icons, no visual hierarchy, no styling to guide the eye. You give a participant a task, something like "find where you'd change your billing plan," and watch which path they take through the hierarchy to answer it.
Because participants have only labels and structure, a tree test isolates the information architecture. If they cannot find the correct node, visual details such as button size or color are not responsible. The item's category or label may not match how participants think about the task.
That makes tree testing well suited to a specific moment in a project: after you have a candidate structure but before you've invested in visual design around it, so you're not paying to redo layouts because the underlying navigation was wrong to begin with.
Tree testing vs. card sorting
Tree testing and card sorting get grouped together often, and they do pair well, but they answer different questions at different points in the process.
Card sorting is generative. You hand participants a set of cards, usually the individual pages or features that need to live somewhere, and ask them to group those cards into categories that make sense to them (in an open sort, they can also invent their own category names). The output is raw material you use to draft a structure.
Tree testing is evaluative. You've already built a structure, and now you're checking whether it works by asking people to navigate it to complete specific tasks. Participants only move through the hierarchy you built, picking a node at each level until they land on an answer or give up.
| Card sorting | Tree testing | |
|---|---|---|
| Purpose | Generate a structure | Evaluate a structure |
| What participants do | Group cards into categories | Navigate a tree to complete tasks |
| Best timing | Before you've drafted an IA | After you've drafted an IA, before visual design |
| What it reveals | How people naturally categorize your content | Whether people can find specific things in what you built |
The two methods often run in sequence: a card sort generates candidate structures, then a tree test evaluates the selected candidate before visual design begins. Card sorting: a practical guide explains the generative part of that sequence.
Designing tree test tasks
Task wording strongly affects a tree test, so review it carefully before building the tree.
A good task describes a goal in the participant's own terms, not your navigation's terms. "Find where you'd change your billing plan" works because it describes something a person wants to do. "Find the Billing section" doesn't, because it hands them the answer by naming the label you're trying to test.
A few practical guidelines:
- Write tasks around real goals, not menu items. If your task wording reuses a label from the tree, you're testing whether people can read, not whether the label makes sense.
- Keep tasks specific enough to have one right answer, or a small set of them. A vague task produces vague navigation, and you won't be able to tell whether a wrong answer reflects a real problem or an ambiguous instruction. Versive tree test tasks support one or more correct answers, so mark every acceptable destination when the task has more than one.
- Cover the categories you're most unsure about. Don't spend all your tasks confirming a structure you're already confident in. Spend them on the sections you rewrote, merged, or renamed.
- Test one thing per task. A task that requires navigating two unrelated parts of the tree makes it hard to tell which part actually failed.
- Plan for 5 to 10 tasks per session. Enough to cover your riskiest areas, few enough that participants don't fatigue and start guessing.
Order matters too: later tasks benefit from a familiarity with the tree that the first task didn't have. Randomizing task order across participants spreads that learning effect out instead of concentrating it on whichever task happens to be last.
Building the tree
The tree itself should mirror the real structure you're testing, or a serious candidate for it, not a simplified stand-in. Leaving out sections to keep the tree tidy also removes the plausible wrong turns participants would take in the real product, which quietly inflates your success rates.
A few things to get right when building it:
- Match real depth. If your actual navigation is four levels deep in places, test four levels deep. Flattening it hides problems that only show up once people have to keep drilling down.
- Use real labels, not placeholder or internal-only names. The label a participant sees in the test should be the exact label they'd see in the product.
- Include distractor categories. If every task has one unmistakable path, the test provides little evidence about competing labels. Include categories that could plausibly, but incorrectly, hold an answer, as a real architecture would.
- Don't over-trim to fit the test. If a section is part of the structure, leave it in, even if it makes the tree longer to build.
Reading the results: success rate and directness
Two metrics do most of the interpretive work in a tree test, and they answer different questions.
Success rate is the simplest read: what share of participants landed on a correct answer for a given task, versus a wrong answer or giving up entirely. It tells you whether a task, as a whole, works. A task with a low success rate needs attention somewhere in its path.
Directness tells you where in that path things went wrong. A participant who goes straight to the correct node without doubling back is direct; one who picks a category, backtracks, tries another, and eventually lands on the right (or wrong) answer is not. Low success with high directness usually means participants made a fast, confident, wrong choice, which points to a mislabeled or miscategorized node. Low success with low directness points to a structure that's confusing throughout the path, not just at one decision point.
Read the two metrics together rather than treating every failed task alike. If participants confidently choose the same wrong category, revise the label or parent category. If they wander and give up, the structure may be too deep or ambiguous across several levels.
It's also worth looking at the actual paths participants took, not just whether they succeeded. A handful of participants converging on the same wrong node is a much stronger signal than an even scatter of different wrong answers, even if the success rate is identical in both cases.
Running a tree test in Versive
Versive includes Tree Test as a question type, so a tree test can run as its own study or sit inside a larger session alongside surveys, ratings, or AI interview questions.
To build the tree, work in visual mode, dragging and dropping nodes to construct and rearrange the hierarchy directly, or in text mode, typing or pasting the tree as indented text. Text mode is the faster path if you're starting from an existing site map or spreadsheet. Trees can also be imported and reused across studies, so a structure you validate once doesn't need to be rebuilt for the next round of testing.
Each task gets a description and one or more correct answers, matching the task design guidance above: write the description around a real goal, and mark every node that would count as a legitimate answer. You can randomize task order across participants, and configure whether they're allowed to skip a task, go back to a previous step, or select a parent node as their final answer rather than drilling all the way down. Tree depth defaults to 6 visible levels, which covers most real navigation structures without extra configuration.
Results land in a dedicated dashboard: Sankey diagrams showing how participants moved through the tree, including where they backtracked; path analysis breaking down the most common routes per task; success rates split into correct, incorrect, and gave-up; and per-task analytics showing which parts of the structure need rebuilding. Success rates come straight out of the dashboard, and the Sankey diagrams and path analysis are where you read the backtracking behavior that separates a direct path from a wandering one.
Because Tree Test is a question type like any other, it composes with the rest of a study: open with a card sort to generate structure candidates, follow with a tree test on the resulting hierarchy, and close with an AI Question asking participants why they picked the path they did, all in one session. See the question types reference for how tree test settings sit alongside every other type in the builder.
When to run a tree test
Run a tree test after drafting a structure and before building its visual design. At that stage, you can correct labeling and categorization problems before they affect layouts or ship in a redesign.
Pair it with a card sort when you need to generate the structure. Once visual design is in place, follow with a broader usability test to confirm that the structure works with real content. User interviews vs. surveys compares structured and conversational methods.
Add a Tree Test question to a study, build or import the hierarchy, and start with two or three tasks for the sections you are least confident about.
Frequently asked questions
What is tree testing used for?
Tree testing checks whether people can find things in your navigation structure, using a text-only version of the hierarchy with no visual design to help or hide problems.
What is the difference between tree testing and card sorting?
Card sorting is generative: participants group items to help you build a structure. Tree testing is evaluative: participants navigate a structure you already built to see if it works.
What is a good success rate for a tree test task?
There is no universal number, since it depends on the task and the stakes of getting it wrong, but a task where most participants fail or give up is a strong signal that node or label needs to change.
What is directness in a tree test?
Directness measures whether a participant went straight to the correct node or backtracked along the way. A low success rate with high directness usually means a clear but wrong choice, while backtracking points to a confusing structure.
Can I combine tree testing with other question types in the same study?
Yes. A tree test is a question type in Versive, so you can place it alongside surveys, AI interview questions, or other tasks in a single study.
Full reference
Question types
Keep reading
Card sorting: a practical guide
A practical guide to card sorting for UX research: open vs. closed sorts, card and participant counts, and how to analyze the results.
Concept testing methods, compared
Concept testing methods compared: monadic vs. sequential monadic vs. comparative, qualitative vs. quantitative, and how to run each.
Moderated vs. unmoderated usability testing
Moderated vs. unmoderated usability testing: the classic trade-offs, and how AI moderation now blurs the line between them.
