AGENTIC MULTIMODAL AI

Real-Time AI Co-Agent for Responsive Design Learning Environments

GEMLA-Agent is a cloud-native, multimodal agentic AI system designed to collaborate with teachers during live instruction— detecting instructionally significant moments and delivering context-aware instructional support in real time.

Real-time Multimodal Understanding
Pedagogically Grounded
Teacher in Control
Human–AI Collaboration
LIVE
Grade 8 Mathematics
Algebraic Expressions
GEMLA-Agent

Multimodal Input Streams

Teacher Speech
Student Voices
Gestures & Gaze
Written Work
Artifacts

DMAMG Memory Graph

Evolving Classroom Context

Pedagogical Reasoning

!
Instructional Moment
Conceptual connection detected
?
Learning Opportunity
Student reasoning emerging
Participation
Equity pattern identified
Priority
Intervention ranked

Teacher Support

Suggested Teacher Action

Ask a probing question that connects the student's current reasoning to the class discussion.

Real-time instructional prompts
Suggested teacher questions
Alerts about overlooked student ideas

Built for Real-Time Teaching

GEMLA-Agent supports teachers during the moments when instructional decisions matter most.

Grounded in Research

The pedagogical reasoning engine is grounded in Responsive Design Pedagogy.

Multimodal by Design

Speech, gestures, gaze, written work, artifacts and participation patterns.

Open Research Contribution

GEMLA-Agent contributes open-source components, benchmarks and multimodal classroom datasets.

From Classroom Signals to Instructional Action

GEMLA-Agent continuously perceives classroom events, builds contextual memory, reasons pedagogically, coordinates agents and provides actionable teacher support.

Six Computational Layers

1
Multimodal Classroom Data Real-time audio, video, and affective student discourse, gestures, gaze, body movement, collaborative interactions, written mathematical work, diagrams, and manipulatives. Multi-Edge-first processing
2
AI Processing Engine Specialized perception agents transform multimodal signals into structured semantic events using multimodal LLMs, computer vision, and cross-modal fusion. Reasoning optimiser, conceptual understanding, instruction
3
Pedagogical Intelligence Responsive Design Pedagogy grounds the reasoning engine in detecting significant moments, identifying conceptual connections, monitoring participation and equity, and prioritizing interventions. Students, situations, actions
4
Multi-Agent Coordination Distributed agents share memory, process candidate interventions, evaluate instructional utility, negotiate priorities, and dynamically adapt as student situations evolve. Continuous priority rebalancing
5
Teacher Support Interface Low-latency, human-interpretable actions: instructional prompts, suggested questions, noticed idea alerts, grouping and discussion recommendations, and participation equity notifications. Designed for action in heated sessions
6
DMAMG – Shared Temporal Memory Nodes encode multimodal instructional events and inferred conceptual states; temporal edges connect events across episodes for long-horizon reasoning, uncertainty propagation, memory compression, salience estimation, and adaptive retrieval. Persistent situational awareness across agents

Dynamic Multimodal Agent Memory Graph

  • Continuously evolving temporal memory
  • Nodes represent multimodal instructional events
  • Temporal edges capture relationships across events
  • Supports long-horizon contextual reasoning
  • Uncertainty-aware adaptive retrieval
  • Memory compression and event salience estimation

A Distributed Architecture for Human–AI Co-Reasoning

Event-driven multimodal agents continuously update shared contextual memory and adapt to emerging classroom events.

Capture

Audio
Video
Artifacts

Interpret

Speech
Vision
Fusion

Remember

DMAMG
Memory
Context

Reason

Pedagogy
Priority
Equity

Act

Prompts
Questions
Recommendations

Measuring Real-Time Human–AI Co-Reasoning

An open benchmark suite for evaluating multimodal agent systems under realistic streaming conditions.

Technical Performance

Response latency targets below 2–3 seconds, multimodal interpretation accuracy, recommendation precision and recall, and streaming reliability.

Human–AI Collaboration

Teacher trust and adoption, recommendation usefulness, recommendation utilization and perceived instructional support.

Educational Impact

Teacher responsiveness, participation equity, student engagement and comparative student learning outcomes.

From Prototype to Classroom Evaluation

01

Core System Development

Months 1–6
  • Multimodal data pipelines
  • Multi-agent orchestration
  • Pedagogical reasoning modules
  • Simulated classroom testing
02

Pilot Deployment & Dataset Development

Months 7–12
  • Deployment in 10–20 mathematics classrooms
  • Authentic multimodal interaction data
  • Latency and usability evaluation
  • Refinement of intervention policies
03

Comparative Evaluation & Dissemination

Months 13–18

Potential Impact

GEMLA-Agent evaluates whether real-time human–AI collaboration can strengthen teacher responsiveness and student learning.

Student Profiles & Class Rack

Academic history, behavioural patterns, and progress tracking, organised by class for instant access.

Live Teacher Insights

Real-time visibility into classroom dynamics as they happen, so teachers can adjust on the spot.

Career Advisory

Reads strengths, interests, and performance to surface personalised career pathways, starting early.

Teacher Professional Development

Continuous feedback identifies teaching gaps and recommends targeted development resources.

Pedagogical Intelligence

Classroom-level analytics — from attentiveness to physical proximity — beyond what test scores show.

School-Wide Analytics

A bird's eye view of performance, attendance, engagement, and outcomes across classes and grades.

Teacher Responsiveness

Support teachers in responding to student mathematical thinking during live instruction.

Participation Equity

Identify participation patterns and opportunities for more equitable classroom interaction.

Student Engagement

Examine student engagement patterns within multimodal classroom interactions.

Learning Outcomes

Compare learning outcomes between classrooms using GEMLA-Agent and classrooms operating without it.

Reusable Infrastructure

Contribute open-source components, datasets and evaluation frameworks to the AI research community.

Scalable Agentic AI

Explore reusable human–AI co-reasoning infrastructure for dynamic real-world environments.

The people building GEMLA

GEMLA stands for GenAI-Enhanced Multimodal Learning Analytics. A team combining expertise in AI, education, and software engineering to rethink how schools understand learning.

Denish Akuom

Denish Akuom, PhD

Founder and CEO

Innovation leadership, mathematics education, maker-centered learning, responsive pedagogy research, STEM technology innovation, and product strategy.

Irene Jerry

Irene Jerry, MSc

Data and Impact Lead

Data systems, AI and ML applications, analytics, educational technology, monitoring and evaluation, and impact measurement.

Dan Eldon Onguka

Dan Eldon Onguka

Lead Software Engineer

Software development, AWS cloud architecture, AI initiatives including GEMLA, cybersecurity, digital platforms, LMS development, automation, and mobile applications.

Dancun Ochieng

Dancun Ochieng

Program Director and STEM Instructor

Telecommunications engineering, technical systems management, AI applications, and hardware and software integration.

Dorphine Oyugi

Dorphine Oyugi

STEM Instructor

Human-centered learning design, inclusive technology adoption, and problem-solving methodologies.

Elias Wanga

Elias Wanga

Digital Learning and Web Development Coordinator

Website and digital platform development, educational technology resources, online content management, web development instruction, and multimedia production.

From the team

We write about AI in education, what we are learning, and what we are building. First posts coming soon.

Coming Soon
Research

What multimodal classroom AI actually needs to work

A look at the real challenges in building AI that operates during live instruction, and how we approached them in GEMLA.

Coming Soon
Architecture

Why we built the DMAMG: the memory problem in real-time AI

Real-time AI forgets. Here is how the Dynamic Multimodal Agent Memory Graph keeps a lesson in context from the first minute to the last.

Coming Soon
Education

Rethinking how we measure participation in the classroom

Hand raises and verbal answers tell only part of the story. What a fuller picture of participation looks like, and why it matters for equity.

We are getting our first posts ready. Get in touch if you want to be notified when we publish.

Building the Future of Human–AI Co-Reasoning

GEMLA-Agent combines multimodal perception, pedagogical intelligence and cloud-native multi-agent orchestration to support teachers in real time.

Connect With the GEMLA-Agent Team →