Text, voice, or video: choosing an interview mode
Choose an interview mode in Versive, with text and voice recommended for most studies and text only reserved for studies that should not collect recordings.
Every Versive study lets you choose whether participants respond by typing, speaking, or appearing on camera. For most studies, start with text and voice so each participant can choose per question and spoken answers are available when useful. Choose text only when you explicitly do not want to collect recordings or prompt for microphone access.
Available modes
Versive's conversation mode setting depends on which interview engine your study uses.
| Mode | Engine | What participants do |
|---|---|---|
| Text only | Regular AI | Type every answer. |
| Text and voice | Regular AI | Type or speak each answer, question by question. |
| Voice with preview | Regular AI | Speak each answer, with a preview before it's submitted. |
| Voice only | Regular AI | Speak every answer (typing disabled), without a transcription preview. |
| Video | Regular AI | Answer on camera, question by question. |
| Voice to voice | Realtime AI (beta) | Have a live, spoken conversation with the AI interviewer. |
All Regular AI modes proceed one question at a time, whether the participant types, talks, or appears on camera. Voice-to-voice runs on the Realtime AI engine, currently in beta. The AI interviewer speaks questions aloud, responds immediately, and can handle interruptions and back-and-forth conversation. Realtime text is being deprecated and is not recommended for new studies.
When to use text only
Use text only when the study should not collect audio recordings or request microphone access. This can be appropriate for a policy-sensitive study, a participant group with limited connectivity, or research where written editing is part of the intended response.
Some participants may also prefer typing for sensitive topics. They can edit an answer before submitting it, and the AI interviewer can still ask adaptive follow-ups. If recordings are acceptable, text and voice preserves this typing option while also letting other participants speak. See AI follow-up questions: probing without a moderator for details.
What voice captures
People generally talk faster than they type, so voice modes tend to produce longer, less-edited answers. Recordings also preserve tone, hesitation, and emphasis. These details may make text and voice, voice with preview, voice only, or voice to voice a better fit than text only.
Question media, such as an image, Figma prototype, or video, can be attached to most questions in any mode. In a spoken mode, participants can describe what they see as they react to it. This can be useful for research on visual designs.
The AI interviewer can rephrase the next planned question based on the participant's previous answer. In voice mode, rephrasing can make transitions sound less scripted. If exact wording matters, such as for a standardized instrument, turn rephrasing off. See The AI interviewer for the full set of personality and follow-up controls.
For any voice mode, add likely product names, domain terms, and jargon under the interviewer's transcription keywords setting. More accurate transcripts are easier to analyze.
Voice with preview vs. voice to voice
These two are easy to conflate. Voice with preview still runs one question at a time on the Regular AI engine: the participant speaks, sees a preview of what they said, and moves on. Voice to voice runs on the Realtime AI engine and is a continuous, live conversation rather than a series of discrete turns.
If you use voice to voice, the silence threshold controls how long the AI waits after a participant stops talking before responding. The range is 0.5 to 5 seconds, with a default of 1.5. A short threshold may interrupt someone who pauses to think; a long threshold adds delay. Adjust it for the expected speaking pace of your participants.
When to use video
Video mode records participants on camera as they answer, adding visible reactions and body language to their spoken responses. That context can help with product and usability research. On desktop studies, you can pair video with screen recording. Screen recording can capture the browser tab while a stimulus (an image, prototype, or embedded website) is on screen, or capture the full screen for the whole session.
Video adds the most participation requirements because it needs camera access and comfort being recorded. It may suit a smaller, more engaged, or compensated participant pool better than a broad public link.
Accessibility and comfort
Text and voice is the recommended default for most studies because each participant can choose per question. Participants who find typing slow or difficult can speak, while those in a noisy environment or those who prefer typing can use text.
In a text-only study, Show voice tooltip can still invite spoken answers. Leave it off when the reason for choosing text only is to avoid recordings. Video requires a camera throughout the interview, so use it only when visible reactions support the research goal.
Matching mode to your research goal
Use these as starting points:
- Most studies, including broad links, qualitative feedback, discovery, and pulse checks: text and voice, so participants can choose what is comfortable.
- Studies that should not collect recordings: text only on the Regular AI engine, with Show voice tooltip disabled.
- Depth interviews where a reviewable, turn-based transcript matters: voice with preview.
- Live conversations or interviews built around a prototype or image: voice to voice on the Realtime AI engine. It uses two credits per completion, compared with one for Regular AI.
- Usability and product studies where visible reactions matter: video, paired with screen recording on desktop.
Where to set it
Conversation mode lives in Settings → General, under the interview engine setting that determines which modes are available. The silence threshold for voice to voice is in the same section. The AI Interviewer tab contains rephrasing controls and transcription keywords for spoken answers. See Study settings for the full reference.
For the rest of the setup, see Create your first AI-moderated study and Survey question types, and when to use each. You can change the mode later, but changing it after collecting responses can make the participant experience less consistent across the study.
Frequently asked questions
Is voice-to-voice the same as text and voice mode?
No. Text and voice runs on the Regular AI engine and lets participants type or speak each answer in turn. Voice to voice runs on the Realtime AI engine and is a live, spoken back-and-forth conversation.
Can participants still speak their answers in a text-only study?
Yes, if Show voice tooltip is enabled. If you chose text only specifically to avoid collecting recordings, leave that setting off.
Does adding voice or video cost more than a text study?
The interview engine, not the conversation mode, sets the credit cost. Regular AI is one credit per completed interview, and Realtime AI, which powers voice to voice, is two credits per completion.
Full reference
The AI interviewer
Keep reading
Create a study with AI in 60 seconds
See how Versive works as an AI survey generator, turning a plain-language prompt into a full study you can refine and share.
Create your first AI-moderated study
How to run an AI-moderated interview in Versive: create a study, add questions, configure the AI interviewer, and share it.
