Why the questions are designed the way they are
Every question the Acai voice agent asks a student is grounded in decades of research on how genuine learning can be verified through conversation. This isn’t a set of arbitrary prompts — it’s a framework built on cognitive science, educational psychology, and institutional best practice from universities that have been running structured oral assessments at scale for years. Understanding that foundation helps explain why the voice agent asks what it asks, and why oral conversation is one of the most reliable ways to confirm that a student actually did the learning behind their work.
Why conversation reveals what detection tools cannot
Since generative AI became widely available, the higher education sector has converged on a clear conclusion: redesigning assessment works better than trying to detect misconduct after the fact. Cornell University’s Center for Teaching Innovation explicitly does not recommend AI detection algorithms, arguing that trusting relationships and authentic assessment design outperform policing. TEQSA, Australia’s higher education regulator, reached the same conclusion — AI detection tools can’t guarantee integrity, and structural redesign is the only sustainable response — and required all 203 Australian providers to submit institutional action plans by July 2024.
Oral assessment sits at the center of that redesign because of a simple property: spoken responses can’t be pre-fabricated or outsourced in real time. An examiner (or a voice agent) can follow an unexpected thread, ask for elaboration, or probe a connection — and only someone who actually did the work can follow along. As Western Sydney University’s assessment guidance puts it, “there are no standard responses that students can simply reproduce to ensure success.”
This isn’t just intuitive — it’s measurable. A 2023 systematic review and meta-analysis in BMC Medical Education found that structured oral examinations achieve substantially higher reliability and validity than unstructured ones, with inter-rater reliability coefficients reaching 0.78–0.91 when standardized questions and calibrated rubrics are used, compared to far lower consistency in traditional, unstructured vivas. A randomized, double-blind comparative study of medical students reached the same conclusion. Educational psychology research adds the other half of the picture: oral questioning is particularly effective at separating deep learning from surface learning, because students who only memorized material can often recite definitions but struggle the moment they’re asked “why” or “how does this connect” to something else.
The five bodies of research behind every question
The voice agent’s questions are built on five interlocking theories, each contributing a specific design principle.
Bloom’s Revised Taxonomy (Anderson & Krathwohl, 2001) provides the cognitive scaffolding — Remember, Understand, Apply, Analyze, Evaluate, Create. Most exam questions in education have historically clustered at the lowest level; a well-designed oral conversation deliberately weights questions toward the higher-order levels, where surface learning breaks down.
Deep versus surface learning and the SOLO Taxonomy (Biggs & Collis, 1982) supply the quality lens. The key distinction is between multi-structural answers — a list of facts with no integration — and relational answers, where a student links concepts into a coherent whole. That shift from multi-structural to relational is the qualitative leap from surface to deep learning, and it’s exactly what follow-up probing is designed to reveal.
Constructive alignment (Biggs, 1996) requires that oral questions activate the same cognitive verbs as the assignment’s learning outcomes — if a brief asked students to “critically evaluate,” the conversation needs to require evaluation, not recall.
Metacognition and self-regulated learning (Flavell, 1979; Zimmerman, 2002; Schraw & Dennison, 1994) is where the framework gets its sharpest authenticity signal. Genuine metacognitive awareness — knowing how you know something, and what you did when you got stuck — is deeply personal and extremely difficult to fabricate convincingly.
Experiential and constructivist learning (Kolb, 1984; Piaget; Vygotsky) underpins questions about how prior knowledge shaped the work. Piaget’s concepts of assimilation and accommodation predict that authentic learners can identify specific moments of cognitive conflict and resolution — moments that are nearly impossible to invent convincingly after the fact.
What this looks like in a conversation
The voice agent draws on a bank of content-independent question types, each mapped to the research above. A few representative examples:
- The genesis narrative (Remember/Understand) — asking a student to walk through how they approached the work from the very beginning. Genuine authors have a specific story with false starts and pivots; a fabricated account tends to be generic.
- Conceptual ownership (Understand/Analyze) — asking a student to explain a concept or theory they used in their own words. This forces the translation from academic language to personal understanding that only comes with real comprehension.
- Surprise and struggle (Analyze/Evaluate) — asking what challenged or changed the student’s thinking during the process. This draws directly on Piaget’s idea of productive cognitive conflict, and is one of the strongest authenticity signals in the framework because a believable account of genuine intellectual surprise is very hard to invent.
- Transfer to new ground (Apply/Evaluate) — asking the student to apply their own argument to a new scenario, created live in the conversation. Deep learners can extend their reasoning; surface learners stay locked to the examples they memorized.
- Metacognitive strategy awareness (Evaluate/Create) — asking how the student knew when they understood something well enough, and what they did when they got stuck. This is considered the single most powerful authenticity indicator, because it requires lived experience of struggling with and resolving a real problem.
Follow-up probing matters more than the initial question itself: it’s the ability to pursue an unexpected thread that separates a real conversation from a static Q&A, and it’s exactly where fabricated answers tend to fall apart.
Validated in real institutions, not just theory
This isn’t a purely academic exercise — versions of this approach are already running at scale. The Interactive Oral Assessment (IOA) model, developed jointly by Dublin City University and Griffith University, has been scaled to classes of 750 students at 10–15 minutes per student. The University of Melbourne classifies IOAs as a form of secure assessment, and the University of Sydney places oral assessment in its “secure assessments” lane for verifying individual achievement. TEQSA’s own 2025 guidance on assessment reform in the age of AI incorporates oral assessment components across all of the institutional models it recommends.
The International Baccalaureate’s viva voce guidance frames the same conversation as a “celebration of the completion of the essay” — and that’s the spirit the voice agent is designed to carry into every conversation. It isn’t there to catch students out. It’s there to give students who genuinely did the work a structured, low-pressure way to make that work visible — because the questions that reveal authentic learning are, by design, the same questions that let a genuine learner show what they know.
Sources
Academic integrity & assessment reform
- Cornell University Center for Teaching Innovation — AI and academic integrity
- TEQSA guidance on assessment reform in the age of AI (via Instructure)
- Western Sydney University — Designing and Assessing Vivas
Reliability & validity of structured oral assessment
- Abuzied & Nabag (2023), “Structured viva validity, reliability, and acceptability as an assessment tool in health professions education: a systematic review and meta-analysis,” BMC Medical Education
- Imran, Doshi & Kharadi (2019), “Structured and unstructured viva voce assessment: A double-blind, randomized, comparative evaluation of medical students,” PubMed
- Distinguishing deep from surface learning through oral questioning — PubMed Central
- Oral responses cannot be pre-fabricated in real time — PubMed Central
Cognitive & learning theory
- Bloom’s Revised Taxonomy — University of Illinois Chicago teaching guide
- SOLO Taxonomy overview — Pam Hook
- SOLO Taxonomy levels of response quality — ERIC
- The qualitative shift from surface to deep learning — Teaching Times
- Metacognition: knowledge and regulation of cognition — Wikipedia
- Flavell’s taxonomy of metacognitive knowledge
- Zimmerman’s three-phase model of self-regulated learning — ERIC
- Schraw & Dennison’s metacognitive awareness inventory — ResearchGate
- Kolb’s experiential learning theory
- Piagetian assimilation and accommodation — PubMed Central
Institutional models & practice guidance
- Dublin City University — Interactive Oral Assessment
- Macquarie University TECHE — Griffith University’s Interactive Oral Assessment at scale
- University of Melbourne CSHE — Interactive Oral Assessments guidance
- University of Sydney — AI and assessment, secure assessment lanes
- International Baccalaureate viva voce guidance — “celebration of learning” framing