Perception
Reads language, images, audio, video, or environment state.
Agent AI survey
A classroom reading of Agent AI: Surveying the Horizons of Multimodal Interaction
Reads language, images, audio, video, or environment state.
Changes a physical or digital environment.
Observes the result and adjusts the next step.
Agent AI participates in a task loop; it does more than return advice.
2024 paper framing
Use the January 2024 survey as conceptual vocabulary—not as a 2026 catalog of current systems.
A building block, not the entire agent.
Language plus sensory and contextual signals.
Behavior occurs inside a physical or virtual environment.
Measure outcomes, feedback, and risk—not only text quality.
Historical framing: durable concepts, dated examples.
Concept correction
The difference is the task loop—not simply the number of media types a model can read.
Understands different information sources: text, image, video, audio, or screen state.
Uses perception plus state, goal, action, and feedback to decide what happens next.
More media is not the same as more agency.
Hypothetical classroom test
Treat “students look confused” as an uncertain interpretation—not as a reliable fact about students.
A system processes a hypothetical classroom signal.
It predicts possible confusion—with uncertainty.
It recommends a response and observes feedback.
Class activity only: do not record or upload real student audio or video.
Reusable agent test
Use the same checklist for a robot, game agent, classroom tool, or software assistant.
What can it observe?
Where is it grounded?
What can it change?
What happens next?
What proves progress?
Who could be harmed?
Capability and authority are separate design decisions.
Ethics stakes
Signals such as posture, accent, eye movement, lighting, or camera quality are ambiguous—not evidence of intent.
Screen, face, voice, grade, or classroom behavior.
A probabilistic and potentially biased inference.
Reteach, flag, report, restrict, or change a learning path.
Higher agency raises the cost of a mistaken inference.
Learning boundary
The boundary should be observable student understanding—not merely polished final output.
The student makes decisions, checks evidence, and can defend the work.
The system supplies the reasoning while the student submits the surface output.
Ask: can the student explain, verify, adapt, and defend the result?
Operational policy
A usable policy translates ethical language into observable permissions, duties, and boundaries.
Which systems are approved?
What help is permitted?
What may not be delegated?
What assistance must be reported?
Operational policy replaces slogans with responsibilities.
Classroom governance
These are discussion principles for classroom-agent design—not a replacement for institutional policy.
People and data are not automatically available to an agent.
Recommendations and actions need a reviewable trail.
High-stakes decisions remain with accountable humans.
Protect course, peer, and restricted information.
An agent may suggest; accountable humans must decide and hear appeals.
Takeaway
The upgrade is AI participating in perception, action, and feedback—inside explicit human boundaries.
A chatbot proposes what a person might do.
An agent perceives, acts, and receives feedback.
A human reviews authority, evidence, risk, and consequences.
Design the authority boundary—not only the capability.