CPTS 541 – Computer Vision

This course page is the primary entry point for course materials. Canvas is used for assignment release/submission, grading, and course announcements.

Course Overview

CPTS 541 – Computer Vision is a graduate-level course on the principles, models, and systems that enable machines to understand and interact with the visual world.

The course follows a six-stage progression:

  • Visual learning essentials: image formation, matching, learning, neural networks, CNNs, and vision transformers
  • Reusable visual representations and pretraining
  • Open-world and promptable perception
  • Visual world representation across time, 3D space, and possible futures
  • Multimodal reasoning, embodied intelligence, and visual action models
  • Reliability, evaluation, and efficient deployment

Rather than treating computer vision as a collection of isolated models, the course studies how visual measurements, data, objectives, representations, architectures, task interfaces, and deployment constraints interact in an end-to-end visual system.

The course combines conceptual understanding, mathematical reasoning, practical implementation, empirical comparison, and a course project involving modern computer vision methods.

Learning Outcomes

By the end of the course, students should be able to:

  • Explain how sensing, sampling, geometry, objectives, and optimization determine the information available to a visual system
  • Implement and analyze core visual learning methods, including local correspondence, classifiers, neural networks, CNNs, transformers, and pretrained encoders
  • Compare major visual pretraining objectives and select an appropriate interface for zero-shot use, probing, fine-tuning, or parameter-efficient adaptation
  • Formulate and evaluate detection, grounding, segmentation, video, three-dimensional, and generative vision tasks
  • Analyze and build multimodal, embodied, and vision-language-action systems in terms of architecture, data, grounding, memory, planning, action representation, and closed-loop behavior
  • Design controlled empirical comparisons and evaluate visual systems under distribution shift and deployment constraints, including robustness, reliability, efficiency, and failure analysis

Syllabus and Schedule

Lecture Units and Slides

The course slides are organized into 13 logical lecture units. Each unit typically contains three classes. The calendar schedule, including holidays and presentation dates, is provided in the course syllabus and Canvas.

Unit Topic Slides
1 [CV] Visual Measurement, Matching, and Learning Slides
2 [NN] Neural Networks, CNNs, and Vision Transformers Slides
3 [Pretrain] Self-Supervised Visual Pretraining Slides
4 [VLM] Vision-Language Pretraining and Adaptation Slides
5 [OD] Object Detection and Visual Grounding Slides
6 [Seg] Segmentation and Promptable Perception Slides
7 [Video] Motion, Tracking, and Video Understanding  
8 [3D] Depth, 3D Vision, and Spatial Representation  
9 [GenAI] Image/Video Generation and World Models  
10 [MM] Multimodal Reasoning  
11 [Embodied] Embodied Intelligence  
12 [VLA] Visual-Language-Action Models  
13 [ReliableAI] Reliability, Evaluation, and Deployment  

Assignment Materials

Assignment Topic/Section Materials
HW1 [CV] + [NN] Visual Evidence and Learned Recognition Starter Code
HW2 [Pretrain] + [VLM] Pretrained Representations and Vision-Language Adaptation Starter Code
HW3 [OD] + [Seg] Detection, Grounding, and Promptable Segmentation Starter Code
HW4 [Video] + [3D] Learning from Time and Space  
HW5 [GenAI] + [MM] Generative Vision and Multimodal Evaluation  
HW6 [Embodied] + [VLA] + [ReliableAI] Embodied VLA and Reliable Visual Control