RobinAI Education Radar

Evaluation Tool · Updated 2026-10-04

How can a shared rubric compare education AI tools?

Written and maintained by Robin · Updated

This answer synthesizes public sources. Check the evidence and limits before applying it. Review the evidence

Link to this answer
Citation text
Robin. How can a shared rubric compare education AI tools?. 2026-10-04.
Compare tools for the same learner group, subject, and task, then record learning fit, age fit, output quality, teacher control, data safety, accessibility, evidence, full cost, and exit readiness. A skills-validation product and gaps in school-tool evaluation add requirements for validity and independent review; we recommend marking missing records as evidence needed.
https://edu.hhhh.life/en/guide/ai-tool-evaluation-rubric/#answer
19direct citations
18source organizations
2026-10-04evidence through

READ THIS FIRST

Three judgments to remember

  1. 01

    Start with learning fit and observable value

  2. 02

    Set separate gates for safety, data, and human control

  3. 03

    Evidence, cost, and exit determine durable usability

ACTION PLAN

What to do

Work through five steps in order, preserving process records and the linked evidence. Return to an earlier step when conditions change.

  1. 01

    Define the problem and boundary

    State the task, age, subject, users, and required human decisions and screen out misaligned products first. Sixteen independent tool reviews add test cases for instructional fit, error handling, and general-chatbot risk.

    View 7 sources for this step
  2. 02

    Record the baseline and owners

    Create nine records marked verified, evidence needed, or unmet; agree data, human-control, and exit stop conditions in advance.

    View 7 sources for this step
  3. 03

    Run a bounded practice

    Test candidates on the same course material and tasks, recording errors, human edits, time, and accessibility.

    View 6 sources for this step
  4. 04

    Check outcomes and costs

    Check original studies, population, sample, comparison, independent work, and delayed outcomes; name the study type and omissions.

    View 6 sources for this step
  5. 05

    Expand, adjust, or exit

    Document item-level results, unresolved questions, approved pilot scope, and exit requirements, with an expiry date.

    View 5 sources for this step

REVIEW CHECKLIST

Check before proceeding

  1. □

    Is evaluate an education AI tool tied to an observable task, covered population, and prohibited-use boundary?

    EU and OECD Release a Primary and Secondary AI Literacy Framework Defining 19 Competencies · Teacher Generative AI Guidelines State That AI Grading Cannot Directly Serve as the Final Evaluation of Open-Ended Work · OECD Warns That Better AI-Assisted Task Performance Does Not Necessarily Mean Real Learning · iFlytek Launches Spark Intelligent Grading Machine M50 with Error Diagnosis · Instruction Partners Releases 16 Deep Dives on AI Learning Tools, Warns of General Chatbot Risks · Florida State Board of Education Approves AI Framework for K-12 and State Colleges · Learning Commons Launches Open Platform for K-12 Edtech
  2. □

    Are owners, human review, data handling, incident reporting, and appeals explicit?

    UK Adds Children's Mental Health and Manipulation Risks to Safety Standards for Educational AI Products · UK Expands Generative AI Data Guidance for Schools, Clarifying Personal Information, Bias, and Supplier Risks · World Digital Education Alliance Releases Two Standards Covering the Full Educational AI Life Cycle and Smart Campuses · 35 New Mexico Lawmakers Urge Education Department to Add Independent Testing and Parental Consent for Amira AI Literacy Tool · Philippine Developer Launches Hiraia, an Offline AI Science Tutor That Runs on $100 Android Phones · Oxford University Press Expands AI Study Assistant to Business, Politics and Science Trove
  3. □

    Do results include independent performance, sustained use, workload, safety events, and group differences?

    OpenAI Releases Learning-Outcome Measurement Tools, Shifting Evaluation Toward Reasoning and Mastery · EduClaw-Bench Places AI Tutors in a Continuous 30-Day Simulated Learning Relationship · Acumen Acquires EduCorePro, Applying AI to University Admissions Review and Fraud Prevention · Pearson to Acquire Workera, Adding AI-Native Skills Assessment to Enterprise Learning · EdReports Review of 10 K-12 Curriculum and EdTech Vendors Finds Only One With Third-Party Evidence That AI Features Improve Learning
  4. □

    Do continue, adjust, pause, and exit decisions each have a threshold, date, and owner?

    EU and OECD Release a Primary and Secondary AI Literacy Framework Defining 19 Competencies · OpenAI Releases Learning-Outcome Measurement Tools, Shifting Evaluation Toward Reasoning and Mastery

CURRENT ANSWER

How we answer today

Each judgment links to the relevant news and original sources. New evidence enters the corresponding dimension.

01

Start with learning fit and observable value

Curriculum frameworks, teacher guidance, and evidence reviews connect features to learning goals and independent performance; grading-device speed and error labels need common-task testing. Sixteen independent tool reviews add test cases for instructional fit, error handling, and general-chatbot risk.

View 7 direct sources
02

Set separate gates for safety, data, and human control

Child-product standards, school data guidance, and lifecycle standards cover admission, monitoring, and accountability. New cases let the rubric compare three conditions: independent validity for high-risk assessment, data minimization in an offline tool, and curriculum grounding in an in-textbook assistant.

View 6 direct sources
03

Evidence, cost, and exit determine durable usability

Outcome tools, longitudinal benchmarks, and product change require method review, durable performance, and migration readiness.

View 5 direct sources
04

Copyable nine-item record and decision rule

Candidate and edition: ____; learner group / subject / common task: ____; review date: ____. Complete 1 learning fit, 2 age and account terms, 3 output accuracy, 4 teacher control, 5 data and safety, 6 accessibility, 7 research evidence, 8 full cost, and 9 exit readiness. For each: “status: ____; document or trial example: ____; gap: ____; reviewer: ____”. Resolve stop items before comparing eligible tools on quality and effort. Mark missing records as “evidence needed”. This editorial rubric has no research-validated weights or thresholds; the school should agree them before testing.

View 3 direct sources

EVIDENCE BOUNDARY

Limits to keep in mind

These limits determine how strong a conclusion the page can support.

  1. 01

    These steps are editorial recommendations informed by public sources. The full workflow has not been validated as an intervention; adapt it to local curricula, age, and resources.

  2. 02

    Product features, coverage, and participation establish an implementation entry; learning effects require independent tasks, delayed measures, and disaggregated results.

RELATED QUESTIONS

What else do readers ask?

Each adjacent search question receives a concise answer linked to its supporting evidence.

01

How can reviewers avoid scoring only the demo?

Give candidates the same course material, learner description, and task. Save original outputs, errors, teacher edits, and time. Add independent no-AI work for learning tasks and test accessibility on the school’s devices and support needs. We suggest retaining failures; a supplier demo does not replace local verification.

View 2 sources for this answer
02

What should be checked in a claim that a tool improves attainment?

Record participants, grade and subject, sample, comparison group, task, duration, independent work, delayed measurement, and funding. Identify whether the outcome is the assisted current task or later independent work. Label randomized trials, observational studies, product self-tests, and simulations separately, preserving limits where the population differs from the school.

View 2 sources for this answer
03

How should a high-scoring tool with an unmet essential condition be handled?

Document the unmet condition and affected use, then request evidence or remediation. Defer the relevant workflow if data use or human review remains unresolved. Continue recording other dimensions and limit approval to verified conditions. This editorial method requires the school to agree stop items and responsible staff in advance.

View 2 sources for this answer

EVIDENCE INDEX

Evidence index

Sorted by public date, preserving only verifiable records and original sources.

View 19 related records
Industry & CompaniesGlobalOriginal publication

Pearson to Acquire Workera, Adding AI-Native Skills Assessment to Enterprise Learning

Pearson announced on September 29, 2026 that it has agreed to acquire Workera, an AI-native skills intelligence platform. Workera combines agentic AI, psychometrics and adaptive assessment to measure demonstrated proficiency through role-specific scenarios and simulations, and its customers include ServiceNow, Accenture and the United States Space Force. The acquisition is expected to close in H2 2026 subject to customary closing conditions and any required regulatory filings or approvals, after which Workera will become part of Pearson's Enterprise Learning & Skills business unit.

TecHRProduct Teams / Schools
ResearchGlobalLead date

EdReports Review of 10 K-12 Curriculum and EdTech Vendors Finds Only One With Third-Party Evidence That AI Features Improve Learning

The nonprofit EdReports released a new brief reviewing AI features from 10 K-12 curriculum and edtech vendors. Only one of the 10 vendors provided third-party evidence that its embedded AI features improved educational outcomes. The brief also states that existing frameworks often analyze privacy and safety of AI integrations, but lack consistent ways to determine whether AI can strengthen learning or preserve curriculum coherence.

EdSurgeSchools / Product Teams
Policy & GovernanceGlobalOriginal publication

Florida State Board of Education Approves AI Framework for K-12 and State Colleges

The Florida State Board of Education last Wednesday approved two AI policy rules. Public school districts and charter school governing boards must amend their internet safety policies to include AI guardrails by July 1, 2027. Schools must notify parents when teachers approve an AI instructional tool, including the platform name, classes, and nature of student interaction, and provide a process for parents to object and alternatives. AI tools used in PreK-5 require additional age-appropriateness reviews, and the 28 state college boards of trustees must adopt AI use and limitation policies.

The Miami TimesFamilies / Educators
Industry & CompaniesGlobalLead date

Learning Commons Launches Open Platform for K-12 Edtech

Learning Commons, backed by the Chan Zuckerberg Initiative, launched an open platform for educators and edtech developers, alongside six new partner organizations. The platform has two main components: the Knowledge Graph and Evaluators. The Knowledge Graph publishes K-12 data on GitHub under an open license, including academic standards from all 50 states, learning progressions and curricula. Evaluators measure whether AI-generated content is accurate, grade-appropriate, evidence-based and aligned to state standards. Since its 2025 launch, its datasets have been downloaded about 20,000 times, with more than 70 partners.

EdSurgeEducators / Product Teams
ResearchGlobalOriginal publication

Instruction Partners Releases 16 Deep Dives on AI Learning Tools, Warns of General Chatbot Risks

Education consulting nonprofit Instruction Partners released a large-scale evaluation of AI-powered learning tools, including 16 deep dives into individual products, covering 20 tools and 16 school systems, with interviews of teachers, students, district and building leaders, and product developers. The analysis found the biggest risks from general-purpose chatbots, which students may use to avoid effortful thinking, while purpose-built instructional tools showed more promise but none was ready to do the pedagogical job independently.

Education WeekSchools / Educators
AI TutoringGlobalOriginal publication

Philippine Developer Launches Hiraia, an Offline AI Science Tutor That Runs on $100 Android Phones

Luis Buenaventura, a member of the Blockchain Council of the Philippines, has developed Hiraia, an open-source AI science tutor that runs fully offline on entry-level Android phones costing around $100, supporting Tagalog, Bisaya, and English. Built on Sea AI Lab's Sailor 2 model and a continued-pretraining fork of Qwen 3.5-2B, the app includes over 40,000 science facts and 30,000 illustrations aligned with the Department of Education's MATATAG curriculum. The project received a 1 million peso research grant from the Tether Foundation and is currently in early alpha (v0.3.1), with the developer seeking academic partners for classroom pilots.

BitPinasStudents / Educators
Policy & GovernanceGlobalOriginal publication

35 New Mexico Lawmakers Urge Education Department to Add Independent Testing and Parental Consent for Amira AI Literacy Tool

A bipartisan group of 35 New Mexico state lawmakers sent a letter on Aug. 25 to Public Education Department Secretary Mariana Padilla, calling for greater transparency and additional independent testing to evaluate how accurately Amira, an AI literacy testing tool, measures students' reading proficiency. The tool has been required statewide for K-2 students since the 2025-26 school year, with students reading aloud to a digital avatar. Lawmakers said Amira has collected thousands of student voice recordings and has access to names, genders, birthdays, locations and other sensitive information, and asked that parental consent be required before children use the tool.

The Taos NewsFamilies / Educators
Learning ToolsGlobalOriginal publication

Oxford University Press Expands AI Study Assistant to Business, Politics and Science Trove

Oxford University Press announced on 9 September 2026 that its AI Study Assistant has expanded to Business, Politics and Science Trove, following its October 2025 launch on Law Trove, so all Trove platforms now offer the tool. It generates summaries, answers and targeted quizzes drawn only from textbook content, and the Science Trove version includes textbook images, figures and their original captions. An AI literacy module covering generative AI fundamentals, academic integrity, critical thinking and prompting is available to all Trove users.

Oxford University PressStudents / Educators
Teacher ToolsChinaOriginal publication

iFlytek Launches Spark Intelligent Grading Machine M50 with Error Diagnosis

On September 1, 2026, iFlytek released the Spark Intelligent Grading Machine M50, which can grade and analyze homework for an entire class within minutes, forming an hour-level personalized teaching loop. Its error cause system includes over 4,000 labels and passed expert appraisal. The device supports 50g low-weight paper with a jam rate of 0.4‰.

iFlytek EducationEducators / Schools
Policy & GovernanceChinaPolicy publication

Teacher Generative AI Guidelines State That AI Grading Cannot Directly Serve as the Final Evaluation of Open-Ended Work

In December 2025, the Expert Steering Committee for Teacher Workforce Development under China's Ministry of Education released the Guidelines for Teachers' Use of Generative Artificial Intelligence (Version 1), covering learning, teaching, student development, evaluation, administration, and research.

Guidelines for Teachers' Use of Generative Artificial IntelligenceSchools / Educators