- Updated: August 13, 2026
- 7 min read
Scaffolding Critical Engagement with GenAI: Transforming Ethnic Minority Preparatory Students’ Collaborative Discourse in Prompt Engineering Tasks
Direct Answer
The paper introduces a design‑based intervention that scaffolds ethnic‑minority preparatory students in China to move from passive reliance on generative AI (GenAI) toward critical, collaborative co‑creation during prompt‑engineering tasks. By embedding teacher‑modeled contrasting cases and a human‑in‑the‑loop workflow, the study demonstrates measurable gains in students’ prompt self‑efficacy and a shift in discourse from “copy‑and‑paste” to strategic evaluation.
Background: Why This Problem Is Hard
Generative AI promises to level the educational playing field by providing instant access to high‑quality explanations, multilingual support, and personalized tutoring. Yet, the very convenience that makes GenAI attractive also creates a cognitive shortcut: students can treat the model as an oracle that delivers answers without requiring deep reasoning. This “answer‑engine” mindset is especially risky for ethnic‑minority learners who already face language barriers, limited exposure to advanced resources, and systemic biases in classroom interaction.
Existing interventions typically focus on technical skill acquisition—teaching students how to phrase prompts or select model parameters. While these tactics improve surface‑level proficiency, they rarely address the epistemic habits that underlie critical engagement. Moreover, most AI‑enhanced curricula assume a homogeneous learner profile, ignoring cultural nuances that shape how minority students negotiate authority, trust, and agency when interacting with an algorithmic partner.
Consequently, educators confront three intertwined challenges:
- Authority bias: Students may accept AI‑generated content as definitive, especially when it appears fluent in a second language.
- Prompt paralysis: Over‑reliance on the model can freeze students’ own ideation, leading them to wait for the AI to “solve” the problem.
- Lack of collaborative discourse: Without structured peer interaction, learners miss opportunities to critique, refine, and co‑construct prompts.
Addressing these issues requires more than a toolbox of prompt‑writing tips; it demands a pedagogical scaffold that repositions the student as an active gatekeeper of AI output.
What the Researchers Propose
The authors present a three‑week, human‑in‑the‑loop scaffolding framework that blends teacher modeling, contrasting case analysis, and collaborative discourse prompts. The core idea is to embed “critical engagement checkpoints” throughout the prompt‑engineering workflow, forcing students to articulate why a prompt is chosen, evaluate the AI’s response, and iteratively improve it with peers.
Key components of the framework include:
- Teacher Modeling with Contrasting Cases: Instructors demonstrate two divergent approaches to the same task—one that relies on blind copying and another that emphasizes hypothesis testing and peer feedback.
- Human‑in‑the‑Loop Loop: Students submit a prompt, receive a GenAI response, and then must annotate the output, flagging inaccuracies or biases before proceeding.
- Collaborative Prompt‑Planning Boards: Small groups co‑design prompts on shared digital canvases, negotiating language choices and evaluation criteria.
- Reflective Journaling: After each session, learners write brief reflections on their decision‑making process, reinforcing metacognitive awareness.
Collectively, these elements aim to transform “strategy talk” from a coordination tool for rapid copying into a metacognitive scaffold for critical evaluation.
How It Works in Practice
During each class, the workflow unfolds in four stages:
1. Prompt Initiation
The teacher presents a real‑world problem (e.g., summarizing a historical event in Mandarin). Students draft an initial prompt individually, then post it to a shared board.
2. AI Generation & Annotation
GenAI returns a draft answer. Students annotate the text, marking factual errors, linguistic ambiguities, or cultural misrepresentations. This step forces them to confront the model’s limitations.
3. Peer Co‑Construction
Groups discuss annotations, negotiate revisions, and collaboratively rewrite the prompt. The revised prompt is fed back to the model, producing a second iteration.
4. Reflective Consolidation
Each learner records a short reflection on what changed, why the new prompt succeeded, and how the peer dialogue shaped the outcome. The teacher then highlights contrasting cases to illustrate effective versus ineffective strategies.
What distinguishes this approach from conventional “teacher‑demo‑student‑practice” cycles is the intentional insertion of a critique loop that treats the AI as a partner rather than a black‑box oracle. By making the evaluation step visible and collaborative, the framework cultivates epistemic agency.

Evaluation & Results
The researchers employed a mixed‑methods evaluation:
- Epistemic Network Analysis (ENA): Quantified shifts in discourse patterns across the three‑week period.
- Thematic Analysis of Reflections: Identified emergent themes such as “authority bias reduction” and “prompt paralysis mitigation.”
- Paired‑samples t‑tests: Measured changes in self‑reported prompt‑engineering self‑efficacy before and after the intervention.
Key findings include:
- Strategic Repurposing: Early sessions showed students using “strategy talk” to coordinate rapid copying. Post‑intervention, the same language was repurposed to scaffold critical evaluation and peer co‑construction.
- Increased Prompt Self‑Efficacy: The average self‑efficacy score rose from 3.2 to 4.6 on a 5‑point Likert scale (p < 0.01), indicating heightened confidence in designing and revising prompts.
- Discourse Realignment: ENA revealed a statistically significant increase in nodes representing “critical questioning” and “peer justification,” while “copy‑and‑paste” nodes declined sharply.
- Qualitative Shifts: Students reported feeling “more like gatekeepers” of AI content, noting that the teacher’s contrasting cases helped them recognize when the model was over‑confident or culturally insensitive.
Collectively, these results demonstrate that targeted scaffolding can overturn the default passive consumption pattern and foster a collaborative, evaluative stance toward GenAI.
Why This Matters for AI Systems and Agents
For AI practitioners building agents that interact with learners, the study offers three actionable insights:
- Design for Human‑in‑the‑Loop Evaluation: Embedding annotation interfaces that surface model errors encourages users to remain critical, reducing the risk of over‑trust.
- Leverage Contrastive Demonstrations: Providing side‑by‑side examples of effective versus ineffective prompt strategies can accelerate the development of epistemic agency in users.
- Facilitate Collaborative Prompt Boards: Multi‑user editing environments, akin to shared whiteboards, enable peer‑driven refinement, which aligns with the “co‑construction” phase of the framework.
These design principles can be operationalized on platforms such as the Workflow automation studio, where developers can script annotation checkpoints and integrate collaborative canvases directly into AI‑driven tutoring bots. By treating the AI as a teammate rather than a teacher, system architects can mitigate cognitive laziness and promote deeper learning outcomes.
What Comes Next
While the study makes a compelling case for scaffolding, several limitations warrant further exploration:
- Scalability: The intervention was conducted with 78 students in a controlled setting. Scaling to larger, more diverse cohorts will require automated scaffolding cues and adaptive feedback mechanisms.
- Cross‑Cultural Validity: The research focused on ethnic‑minority preparatory students in China. Replicating the framework in other linguistic and cultural contexts will test its generalizability.
- Long‑Term Retention: The three‑week window captures immediate gains, but longitudinal studies are needed to assess whether critical engagement persists after the scaffolding is removed.
Future work could integrate advanced multimodal models that not only generate text but also provide visual explanations, thereby enriching the annotation step. Additionally, linking the framework to an Enterprise AI platform by UBOS could enable real‑time analytics on discourse patterns, allowing educators to intervene dynamically.
Potential applications extend beyond K‑12 education. Corporate training programs, language‑learning apps, and citizen‑science platforms could adopt the same human‑in‑the‑loop scaffolding to ensure users remain critical evaluators of AI output.
Conclusion
The research demonstrates that technical proficiency with GenAI is insufficient for equitable education; students must also develop epistemic agency. By weaving teacher‑modeled contrasting cases, collaborative prompt boards, and reflective annotation into a human‑in‑the‑loop workflow, the authors succeeded in turning “strategy talk” from a shortcut into a critical thinking tool. The measurable rise in prompt self‑efficacy and the qualitative shift toward active gatekeeping suggest that such scaffolding can counteract cognitive complacency among ethnic‑minority learners.
For AI developers, educators, and policy makers, the study offers a blueprint for designing systems that nurture critical engagement rather than passive consumption. Embedding these principles into next‑generation AI tutoring agents could be a decisive step toward truly inclusive, agency‑centric learning environments.
Further Reading
For a complete view of the methodology and statistical analysis, consult the original arXiv paper.
Andrii Bidochko
CTO UBOS
Andrii Bidochko is an AI entrepreneur and researcher focused on AI agents, reinforcement learning, and autonomous systems. He writes about the technologies shaping the future of machine intelligence, from frontier models and agent architectures to real-world AI applications.