Robin. How should readers judge the progress, value, and lines of responsibility around AI assessment and feedback?. 2026-10-04.
AI assessment and feedback should track assisted performance, process disclosure, and tool-free transfer together. A Maryland randomized trial and a Harvard classroom experiment provide comparison results, while a skills-validation product and reported evaluation gaps reinforce the need for human judgment. Each result must retain its population, task, and measurement boundary.
https://edu.hhhh.life/en/guide/ai-assessment-feedback/#answer
28direct citations
25source organizations
2026-10-04evidence through
READ THIS FIRST
Three judgments to remember
01
Classroom products now connect personalized feedback and writing-process support to assignments.
02
Regulation and teacher guidance require validation and human responsibility in high-stakes evaluation.
03
Reasoning, mastery, sustained engagement, and independent performance require different measures.
CURRENT ANSWER
How we answer today
Each judgment links to the relevant news and original sources. New evidence enters the corresponding dimension.
01
Formative feedback works best as a draft for teacher confirmation
Google Classroom and Writing Coach support teacher feedback workflows, while an iFlytek device connects grading, error labels, and learning analysis to assignments. Workflows should preserve teacher edits, student revisions, and final rationale, and independently test vendor claims about speed and accuracy. The return to oral examinations adds a process-authentication option whose validity, workload, and accessibility need review.
Exams and high-stakes decisions need a higher evidence threshold
UK assessment regulation restricts independent AI scoring, Ireland's advisory group prioritizes exams and evaluation, and outcome tools are beginning to measure reasoning and mastery. High-stakes use needs validation samples, human review, bias checks, and appeals. A Brazilian writing platform, literacy-assessment concerns, curriculum-text risks, a skills-validation platform, and New South Wales rules add different assessment contexts. High-stakes decisions still require independent validation, process evidence, and human judgment.
Feedback speed and genuine learning need separate measures
Tutor CoPilot offers evidence on teacher support and outcomes, Khanmigo's learning-history experiment observes next-question performance, and StudyFetch evaluates responsible prompting. These results cover teachers, immediate tasks, and process behavior rather than one durable effect.
These limits determine how strong a conclusion the page can support.
01
Product features and company cases show workflow design, while independent evidence on accuracy, bias, and long-term outcomes is often sparse.
02
Assessment criteria vary by subject, age, and task, so one model-accuracy figure cannot summarize educational quality. Needed evidence includes blind-rating agreement, teacher edit rates, revision trails, appeal outcomes, and delayed tests.
RELATED QUESTIONS
What else do readers ask?
Each adjacent search question receives a concise answer linked to its supporting evidence.
01
What can currently be confirmed about AI assessment and feedback?
+
Current public evidence can establish policy, curriculum, program, or product progress. Reach and launch figures should retain their own definitions and remain separate from sustained use and learning outcomes.
Does the available material establish learning outcomes?
+
The available material mainly supports policy, implementation, product, or participation progress. Learning effects still require independent tasks, delayed measures, subgroup results, and reproducible methods.
A University of Maryland randomized trial covered 2,379 undergraduates and 30 instructors in fall 2025. In matched sections of the same course, students offered a course-integrated AI tutor finished about four percentage points lower in final grades, or 0.37 standard deviations, while learning management system participation fell 0.90 standard deviations and page views and active days also declined. About 15% of students offered the tutor used it, and nearly 74% of requests sought information, explanations or answers.
A September 2024 Harvard Gazette report described preliminary findings from a Harvard physics experiment involving 194 students. Students in Physical Sciences 2 experienced two lessons over consecutive weeks, alternating between an instructor-guided active-learning lesson and a custom AI tutor used at home. Preliminary analysis found the AI-tutored group's learning gains were about double those of the classroom group, with higher reported engagement and motivation. The study was led by Harvard lecturers Gregory Kestin and Kelly Miller, and the final research was still pending publication at the time.
Pearson announced on September 29, 2026 that it has agreed to acquire Workera, an AI-native skills intelligence platform. Workera combines agentic AI, psychometrics and adaptive assessment to measure demonstrated proficiency through role-specific scenarios and simulations, and its customers include ServiceNow, Accenture and the United States Space Force. The acquisition is expected to close in H2 2026 subject to customary closing conditions and any required regulatory filings or approvals, after which Workera will become part of Pearson's Enterprise Learning & Skills business unit.
The nonprofit EdReports released a new brief reviewing AI features from 10 K-12 curriculum and edtech vendors. Only one of the 10 vendors provided third-party evidence that its embedded AI features improved educational outcomes. The brief also states that existing frameworks often analyze privacy and safety of AI integrations, but lack consistent ways to determine whether AI can strengthen learning or preserve curriculum coherence.
A Stanford working paper released in June 2026, "The Generative AI Learning Penalty: Evidence from Chinese Secondary Education," uses data from about 27,000 Chinese students in grades seven to 12 to examine how self-directed generative AI use affects cumulative learning. It reports that roughly 80 percent of students began using generative AI between 2023 and 2025, and about 50 percent fully outsourced homework after five months of use. By the end of June 2025, college entrance exam scores fell 18 percent and high school entrance exam scores fell 24 percent, with no corresponding learning benefit found.
The UC Irvine School of Education has launched the National Center for Writing Research to Improve Teaching Effectiveness with Generative AI, known as the WRITE AI Center, through a $10 million federal grant from the U.S. Department of Education's Institute of Education Sciences. Over five years, the center will develop and evaluate evidence-based approaches to using AI in writing instruction and conduct two major studies: a nationwide survey and case study, and an implementation and evaluation study using UCI's PapyrusAI platform. A randomized controlled trial is set to fully kick off in year three, covering more than 1,000 community college students across 60 classes in three states.
Learning Commons, backed by the Chan Zuckerberg Initiative, launched an open platform for educators and edtech developers, alongside six new partner organizations. The platform has two main components: the Knowledge Graph and Evaluators. The Knowledge Graph publishes K-12 data on GitHub under an open license, including academic standards from all 50 states, learning progressions and curricula. Evaluators measure whether AI-generated content is accurate, grade-appropriate, evidence-based and aligned to state standards. Since its 2025 launch, its datasets have been downloaded about 20,000 times, with more than 70 partners.
On September 22, 2026, 2U announced a partnership with CodeSignal to bring more than 400 hands-on, practice-based courses to edX, its global online learning platform, for the first time. The first courses are now live, with the full catalog available on edX by the end of the year. Roughly a third to half of the initial catalog focuses on practical AI skills, including prompt engineering, AI literacy for business roles, ethical AI development, and building AI-powered applications. Skill checks are built into the courses, and CodeSignal's AI tutor Cosmo provides one-on-one guidance and instant feedback.
eSchool News reported on September 18, 2026 that Catherine Hartmann, a professor at the University of Wyoming, redesigned her upper-level humanities course around oral assessment after a 2023 assignment in a history of meditation course came back with a copy-pasted AI prompt left at the top. Students practiced discussion and oral explanation throughout the semester, and the final oral exam became the culmination of that work. At the University of Pennsylvania, mathematics professor Robin Pemantle has students work calculus problems at the board while explaining their reasoning aloud.
On 9 September 2026, UNESCO awarded its 2026 ICT in Education Prize during Digital Learning Week in Paris to Finland's Generation AI and Brazil's Redação Paraná. Generation AI, developed by the University of Eastern Finland, the University of Helsinki, the University of Oulu and other partners, provides classroom activities and tools to help students understand and critically assess AI systems; its materials have been accessed more than 200,000 times across over 50 countries, and workshops and school interventions have involved 1,500 students and teachers. Redação Paraná, developed by the State Secretariat of Education of Paraná in Brazil, combines AI-assisted feedback with teacher involvement to support writing and revision, reaching approximately 850,000 students across 2,000 schools, with teachers trained to integrate its materials. The two projects were selected from more than 100 nominations, and the 2026 prize focused on approaches that use AI while encouraging learners to retain critical thinking, creativity and independent judgement.
A bipartisan group of 35 New Mexico state lawmakers sent a letter on Aug. 25 to Public Education Department Secretary Mariana Padilla, calling for greater transparency and additional independent testing to evaluate how accurately Amira, an AI literacy testing tool, measures students' reading proficiency. The tool has been required statewide for K-2 students since the 2025-26 school year, with students reading aloud to a digital avatar. Lawmakers said Amira has collected thousands of student voice recordings and has access to names, genders, birthdays, locations and other sensitive information, and asked that parental consent be required before children use the tool.
Education Week reports that AI-generated text has entered elementary school classrooms alongside classroom library books and early reading curricula. Tools for educators can rewrite articles to different reading levels or generate decodable text. Jean Gunderson, a Title I reading interventionist in South Dakota, said AI can write stories and comprehension questions in minutes, a task that used to take hours, but she must prompt precisely and sort through all output. Researchers flagged three risks: AI's middle-ground tone, tools struggling to hit requested grade levels, and weak alignment with academic standards and curricula.
Coursera previewed Project Helix, an AI-native skills platform, at its annual FWD customer event on September 8, 2026. The platform aims to help organizations address talent gaps and verify workforce capabilities, drawing on over 30,000 global content partners and instructors from Coursera and Udemy. It will generate adaptive learning paths based on business goals expressed in natural language, with broad availability to enterprise customers expected in the first half of 2027.
Recently, the practice of AI-empowered English listening and speaking teaching in Zhuji, submitted by the Zhuji Education Research Center, was selected for the public list of '2026 Smart Education Excellent Cases'. The case uses AI listening and speaking classrooms, now in regular use in 16 schools, covering 92 classes, 58 English teachers, and over 4,000 students. The project adopts a mechanism of pilot verification, scale-up, and dynamic optimization, and has established a tiered training system.
The University of Glasgow's School of Education has released a free 92-page toolkit to help teachers make deliberate decisions about digital and AI technology in learning. Developed by Mark Peart and colleagues, it guides teachers through examining assumptions, lesson design, and ethical checks, without recommending specific tools.
A study tracking about 27,000 students aged 12-18 in China for 30 months found that after adopting generative AI, homework scores rose by 18% while time per assignment fell from 64 to 45 minutes. However, in monthly closed-book exams without AI, scores dropped by 20% within six months, and high-stakes entrance exam performance also declined. The research was conducted by scholars from Stockholm University and the University of Hong Kong.
The Computing Research Association's Education Committee has issued a white paper calling on universities to rethink how computer science students are taught and assessed in the age of generative AI. It proposes four principles, including treating learning as a process, and suggests alternatives like oral exams and code walkthroughs. The paper cites a University of Illinois facility that proctors over 90,000 exams annually.
On September 1, 2026, iFlytek released the Spark Intelligent Grading Machine M50, which can grade and analyze homework for an entire class within minutes, forming an hour-level personalized teaching loop. Its error cause system includes over 4,000 labels and passed expert appraisal. The device supports 50g low-weight paper with a jam rate of 0.4‰.
The New South Wales government released new rules on 1 September 2026 limiting schools to a maximum of one take-home assessment task worth no more than 15 per cent of the school-based assessment mark, or 7.5 per cent of the total HSC mark. The advice from the NSW Education Standards Authority applies to the Class of 2027 beginning HSC studies in Term 4 this year and to students starting Year 11 in Term 1, 2027. HSC major works and some courses including creative arts, technologies and English Extension 2 are exempt, but schools must still authenticate students' work.
On July 16, Ofqual updated its approach to regulating AI in qualifications, continuing to prohibit AI as the sole scorer while allowing validated supporting and quality-assurance uses.
Khan Academy summarized about 20 Khanmigo product experiments conducted between October 2025 and April 2026. The organization says adding recent practice history and prerequisite skills not yet mastered to responses produced a combined 6.1% increase in the rate of independently answering the next question correctly.
Khan Academy official blogEducators / Product Teams
On May 14, StudyFetch launched Active AI Literacy, scoring each student prompt in real time on quality and responsibility and offering improvement suggestions before it is sent.
On April 7, the Irish government announced an external advisory working group on AI in schools as an ongoing, multi-stakeholder governance mechanism to study AI's effects on teaching, learning, and assessment.
On March 4, OpenAI announced a set of tools for measuring learning outcomes, disclosed early research on Study Mode, and said it planned to continue validation through randomized trials.
On February 19, Google launched AI-suggested feedback in Classroom. Gemini can draft personalized written guidance using a student's assignment, grade level, and focus areas specified by the teacher.
On January 21, Google and Khan Academy announced a partnership to enhance Khan Academy's Writing Coach with Gemini models. The product is focused on guidance and feedback during the writing process.
In December 2025, the Expert Steering Committee for Teacher Workforce Development under China's Ministry of Education released the Guidelines for Teachers' Use of Generative Artificial Intelligence (Version 1), covering learning, teaching, student development, evaluation, administration, and research.
Guidelines for Teachers' Use of Generative Artificial IntelligenceSchools / Educators
In live K–12 mathematics tutoring, Tutor CoPilot suggests guiding questions, hints, and conceptual scaffolds to human tutors. A Stanford research summary reports that the randomized trial involved more than 700 tutors and more than 1,000 students.
Stanford SCALEEducators / Schools
KEEP READING
Continue reading
Enter through an adjacent search question or return to a long-term topic for its full evidence base.