CPTS 541 – Computer Vision
This course page is the primary entry point for course materials. Canvas is used for assignment release/submission, grading, and course announcements.
Course Overview
CPTS 541 – Computer Vision is a graduate-level course on the principles, models, and systems that enable machines to understand and interact with the visual world.
The course follows a six-stage progression:
- Visual learning essentials: image formation, matching, learning, neural networks, CNNs, and vision transformers
- Reusable visual representations and pretraining
- Open-world and promptable perception
- Visual world representation across time, 3D space, and possible futures
- Multimodal reasoning, embodied intelligence, and visual action models
- Reliability, evaluation, and efficient deployment
Rather than treating computer vision as a collection of isolated models, the course studies how visual measurements, data, objectives, representations, architectures, task interfaces, and deployment constraints interact in an end-to-end visual system.
The course combines conceptual understanding, mathematical reasoning, practical implementation, empirical comparison, and a course project involving modern computer vision methods.
Learning Outcomes
By the end of the course, students should be able to:
- Explain how sensing, sampling, geometry, objectives, and optimization determine the information available to a visual system
- Implement and analyze core visual learning methods, including local correspondence, classifiers, neural networks, CNNs, transformers, and pretrained encoders
- Compare major visual pretraining objectives and select an appropriate interface for zero-shot use, probing, fine-tuning, or parameter-efficient adaptation
- Formulate and evaluate detection, grounding, segmentation, video, three-dimensional, and generative vision tasks
- Analyze and build multimodal, embodied, and vision-language-action systems in terms of architecture, data, grounding, memory, planning, action representation, and closed-loop behavior
- Design controlled empirical comparisons and evaluate visual systems under distribution shift and deployment constraints, including robustness, reliability, efficiency, and failure analysis
Syllabus and Schedule
Lecture Units and Slides
The course slides are organized into 13 logical lecture units. Each unit typically contains three classes. The calendar schedule, including holidays and presentation dates, is provided in the course syllabus and Canvas.
| Unit | Topic | Slides |
|---|---|---|
| 1 | [CV] Visual Measurement, Matching, and Learning | Slides |
| 2 | [NN] Neural Networks, CNNs, and Vision Transformers | Slides |
| 3 | [Pretrain] Self-Supervised Visual Pretraining | Slides |
| 4 | [VLM] Vision-Language Pretraining and Adaptation | Slides |
| 5 | [OD] Object Detection and Visual Grounding | Slides |
| 6 | [Seg] Segmentation and Promptable Perception | Slides |
| 7 | [Video] Motion, Tracking, and Video Understanding | |
| 8 | [3D] Depth, 3D Vision, and Spatial Representation | |
| 9 | [GenAI] Image/Video Generation and World Models | |
| 10 | [MM] Multimodal Reasoning | |
| 11 | [Embodied] Embodied Intelligence | |
| 12 | [VLA] Visual-Language-Action Models | |
| 13 | [ReliableAI] Reliability, Evaluation, and Deployment |
Assignment Materials
| Assignment | Topic/Section | Materials |
|---|---|---|
| HW1 | [CV] + [NN] Visual Evidence and Learned Recognition | Starter Code |
| HW2 | [Pretrain] + [VLM] Pretrained Representations and Vision-Language Adaptation | Starter Code |
| HW3 | [OD] + [Seg] Detection, Grounding, and Promptable Segmentation | Starter Code |
| HW4 | [Video] + [3D] Learning from Time and Space | |
| HW5 | [GenAI] + [MM] Generative Vision and Multimodal Evaluation | |
| HW6 | [Embodied] + [VLA] + [ReliableAI] Embodied VLA and Reliable Visual Control |