Kathy Garcia
Portrait of Kathy Garcia

Kathy Garcia.

5th-year Ph.D. student · Johns Hopkins University · kgarci18 [at] jhu.edu

Building human-aligned AI through computational cognitive neuroscience.

I am a 5th-year NSF Fellow pursuing my Ph.D. in Computational Cognitive Science at Johns Hopkins University, where I study how humans and machines understand social interactions in the Computational Cognitive Neuroscience Lab with Professor Leyla Isik.

I measure how people perceive dynamic social scenes, benchmark AI models against human brain and behavior, and use that human data to align video and language models with human social perception. My work has been published at NeurIPS (oral) and ICLR, and featured in The Wall Street Journal.

Previously, I earned my B.S. at Stanford University, worked as a fintech data scientist for Fortune 500 companies, and completed an NIH Postbac Fellowship in affective neuroscience and machine learning at Northwestern University, mentored by Professor Robin Nusslock, Dr. Zachary Anderson, and Dr. Cassandra VanDunk.

🏆 NeurIPS 2026 Oral · first author · top 0.4% of submissions NSF GRFP Fellow ICLR 2025 Best Oral · LatinX in AI @ ICML 2× VSS talks · CCN talk 350+ models benchmarked WSJ · PopSci · Marketplace
🎓 Mentorship — first-gen Latina, happy to chat about research, grad school, or careers!

news

research

Humans read social scenes in a glance; AI still struggles. I use large-scale behavior, fMRI, and deep learning to understand how people perceive dynamic social interactions — and then use that human data to build models that see the social world the way we do.

01 · MEASURE

How do humans organize social scenes?

Tens of thousands of odd-one-out judgments reveal the latent dimensions behind social perception.

02 · BENCHMARK

Where do AI models fall short?

350+ image, video, and language models tested against human brain and behavior.

03 · ALIGN

Can human data close the gap?

A compact behavioral signal steers video and language models toward human social understanding.

MIT Stanford Johns Hopkins

research projects

all publications

  1. Garcia, K., & Isik, L. (2026). Behavioral Geometric Supervision Aligns Video Foundation Models with Human Social Perception. Advances in Neural Information Processing Systems (NeurIPS). Oral (top 0.4% of submissions). arXiv
  2. Garcia, K., & Isik, L. (2026). Human Similarity Judgments Reveal Fine-Grained Representations of Relationships and Actions in Dynamic Social Scenes. PsyArXiv preprint.
  3. Garcia, K., Subramaniam, V., Katz, B., & Cheung, B. (2025). Look, Then Speak: Social Tokens for Grounding LLMs in Visual Interactions. NeurIPS 2025 Workshop on Unifying Representations in Neural Models (UniReps). PDF
  4. Garcia, K., & Isik, L. (2025). Semantic and Social Features Drive Human Grouping of Dynamic, Visual Events in Large-Scale Similarity Judgements. Journal of Vision, 25(9), 2621. DOI
  5. Garcia, K., McMahon, E., Conwell, C., Bonner, M. F., & Isik, L. (2025). Modeling Dynamic Social Vision Highlights Gaps Between Deep Learning and Humans. International Conference on Learning Representations (ICLR). PDF
  6. Garcia, K., Conwell, C., McMahon, E., Bonner, M. F., & Isik, L. (2024). Large-scale Deep Neural Network Benchmarking in Dynamic Social Vision. Journal of Vision, 24(10), 716. DOI
  7. McMahon, E., Conwell, C., Garcia, K., Bonner, M. F., & Isik, L. (2024). Language model prediction of visual cortex responses to dynamic social scenes. Journal of Vision, 24(10), 904. DOI

talks & presentations

  1. Dec 2026NeurIPS 2026 — Oral. Behavioral geometric supervision aligns video foundation models with human social perception.
  2. Sep 30, 2026Invited talk — Stanford NeuroAI Lab (host: Dan Yamins). A Computational Account of Dynamic Social Perception.
    Abstract

    Deep neural networks have become powerful models of human object recognition. Can they also explain how humans perceive people interacting? I will present several studies from my doctoral work to identify a gap in existing models, characterize human social visual representations, and test how behavioral supervision can improve model alignment. First, I benchmark over 350 image, video, and language models against human behavioral ratings and fMRI responses to natural social videos. I find that models that approach the noise ceiling in ventral visual regions leave substantial variance unexplained in the lateral pathway, including superior temporal sulcus. Video models improve predictions in several mid-level lateral regions, while language models given human captions best predict social ratings. These results reveal a challenge for models that aim to explain both the brain and behavior for social vision. Then, in a complementary study, to understand the underlying representations of social perception in humans, we collect approximately 50,000 judgments of 250 social videos and identify dimensions of human social structure that combine relationships and actions and are able to predict representational structure in the human brain and behavior. Finally, I introduce behavioral geometric supervision (BGS), which uses these human similarity judgments to guide the learning of video models by combining local triplet constraints with a global RSA objective, updating fewer than 2% of model parameters. Across all models fine-tuned with BGS, human alignment is improved; particularly in V-JEPA 2.1, this nearly triples the variance explained in held-out human similarity judgments, reaching 78% of the estimated noise ceiling and surpassing the strongest caption embedding baseline. The resulting representations capture behavioral variance beyond caption embeddings, improve the readout of social-affective attributes without explicit attribute supervision, and, surprisingly, are even able generalize to out-of-distribution abstract social interactions. Together, these findings identify a behavioral target and demonstrate an effective path toward models of human social vision.

  3. Aug 18, 2026Invited talk — Cusack Lab.
  4. Aug 2026CCN 2026 — Poster, New York. The latent dimensions supporting dynamic social scene perception.
  5. May 2026VSS 2026 — Poster. Behavior-guided fine-tuning makes vision models more human-like in social perception.
  6. May 2025VSS 2025 — Talk. Semantic and social features drive human grouping of dynamic visual events.
  7. Apr 2025ICLR 2025 — Poster. Modeling dynamic social vision.
  8. Aug 2024CCN 2024 — Contributed talk. Dynamic, social vision highlights gaps between deep learning and humans.
  9. Jul 2024LatinX in AI @ ICML 2024 — Oral (Best Oral Award). Social vision models.
  10. May 2024VSS 2024 — Talk. Large-scale deep neural benchmarking of dynamic social vision.
  11. Apr 2024JHU Cognitive Science brown bag — DNN benchmarking.
  12. 2021SfN Global Connectome — Poster. Predicting dimensional symptoms of psychopathology from fMRI.

honors & awards

  • NeurIPS 2026 Oral (112 of 30,709 submissions; top 0.4%)
  • NSF Graduate Research Fellowship (2024)
  • Best Oral Presentation Award, LatinX in AI Workshop, ICML (2024)
  • Sigma Xi Scientific Research Honor Society (2025)
  • NEI Early Career Scientist Travel Grant (2025)
  • John I. Yellott Travel Award for Vision Science, VSS (2025)
  • FoVea Travel and Networking Award (2024)
  • Kelly Miller Fellowship, Johns Hopkins University (2021–2025)
  • Northwestern Interdepartmental Neuroscience Research Fellowship (2020)
  • Stanford El Centro Latino Acknowledgement (2017)

teaching