Research

The problems I want to solve, where the work lands, and the papers.

  1. 01

    LLMs are memorizing, not generalizing enough

    They lean on memorized patterns instead of transferable, compositional structure. How do we train for genuine generalization?

  2. 02

    LLMs have no cerebellum: no fast loop to act and verify in real time

    Learning happens in slow, offline training runs, and every action waits on the whole model. We need to model the cerebellum: a fast forward model that predicts the next step, carries the routine ones itself, and skips the brain loop entirely, waking the slow model only when its prediction misses. That is what a real-time agent needs, most of all a computer-use agent, where every click waits on a round trip through the big model.

  3. 03

    LLMs cannot learn from a stream without forgetting

    Experience arrives once, in order, and never comes back. Learning from it needs a stronger state store than a context window, and learning without backprop that consolidates what matters into memory as the stream passes, rather than waiting for the next offline training run.

  4. 04

    Safe AI cannot be one authoritarian model: it needs a democratic multi-agent society

    A single large model that holds all the power is a single point of failure. Safety should come from a society of agents: many subagents, each with its own safeguards, and a balance of power strong enough that when a share of the agents break, the rest can repair the society, as human societies do. The levers of self-change stay out of reach: reward systems and training pathways are controlled and closed to the agents, and copying weights is banned. Each agent’s freedom and resources are allocated by the society’s consensus rules and laws.

  5. 05

    LLMs decode by classification, not true generation

    Every token is chosen by scoring it against a vocabulary of hundreds of thousands, and that classifier head is the bottleneck for both generation latency and the long tail, where rare tokens are starved of probability. A human does not compare hundreds of thousands of words before producing each one: the activation drives the actuator directly. Generation should work the same way: a truly generative process that emits the output, not a discriminative one that ranks every candidate at every step.

papers

8

Full list on Google Scholar

academic service

reviewing

  • NeurIPS 2024–2026
  • ICLR 2024–2025
  • ICML 2024–2025

teaching assistant

  • DSCI 552 Machine Learning for Data Science (2023, 2025)
  • CSCI 576 Multimedia Systems Design (2022, 2024, 2025)
  • CSCI 567 Machine Learning (2022, 2024)
  • CSCI 566 Deep Learning and Its Applications (2024)
  • CSCI 544 Applied Natural Language Processing (2023)
Zizhao Huloading加载中