Critique about (Action Localization in Videos )
Action Localization in Videos
Khurram Soomro Center for Research in Computer Vision (CRCV)
University of Central Florida (UCF)
Action Recognition
Diving Lifting
Golf
Swing Bench Walking
Action Localization
1. Action Recognition 2. Action Detection
a. Trimmed Videos i. Spatial
b. Untrimmed Videos i. Temporal ii. Spatio-Temporal
Diving
Lifting
Swing Bench
• Cluttered Background
• Multiple Actors/Actions
• Untrimmed Videos
Basketball Dunk
Salsa Spin
Hand Waving/Clapping/Boxing
Challenges: Action Localization
Applications of Action Localization
• Video Search • Action Retrieval • Multimedia Event Recounting • Video Understanding
Outline
I. Online Localization: Online Localization and Prediction of Actions
II. Unsupervised Localization: Unsupervised Action Discovery and Localization in Videos
Outline
I. Online Localization: Online Localization and Prediction of Actions
II. Unsupervised Localization: Unsupervised Action Discovery and Localization in Videos
Online Localization and Prediction of Actions
Online Localization and Prediction of Actions
Soomro, Khurram, Haroon Idrees, and Mubarak Shah, "Predicting the Where and What of actors and actions through Online Action Localization”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016.
Soomro, Khurram, Haroon Idrees, and Mubarak Shah, "Online Localization and Prediction of Actions and Interactions", IEEE PAMI, 2017 (Under Review).
Comparison: Off line vs. Online
1. Whole video available
2. Entire motion of action 3. Action Recognition
4. Action Detection using all frames
1. Partially observed video
2. Limited motion information 3. Action Prediction
4. Action Detection done frame-by-frame
Off line Action Localization Online Action Localization
Motivation
• Existing approaches localizes after an activity has occurred • Cannot anticipate actions • No timely localization of actions
Action Localization in Videos through Context Walk
Problem Definition
SORTED PREDICTION SCORES
Online Localization and Prediction of Actions
+ Recognition= Action DetectionOff line Action Localization
K IC
K IN
G + Prediction= Action DetectionOnline Action Localization
Problem Definition Online Localization and Prediction of Actions
Idea: Foreground/Background Distinction
✓Human Detection ✓Segmentation
✓Object ✓Motion ✓Action Proposals
✓Pose Estimation (Detecting Body Joints)
Online Localization and Prediction of Actions
Idea: Foreground/Background Distinction
✓Human Detection ✓Segmentation
✓Object ✓Motion ✓Action Proposals
✓Pose Estimation (Detecting Body Joints)
Online Localization and Prediction of Actions
Ø Given an input stream of video frames
Ø Goal: - Ø Predict action in current frame Ø Detect action in current frame
Ø Process in a batch of frames
Proposed Framework Online Localization and Prediction of Actions
Ø Extract Superpixels in each frame
Proposed Framework Online Localization and Prediction of Actions
Ø Pose Estimation
Proposed Framework Online Localization and Prediction of Actions
Framework: Online Action Localization
• Goal:
• Bayes Rule:
• State Transition model:
Pose-based Foreground Likelihood
= Location representing bounding box at time t
Superpixel-based Foreground Likelihood
State- Transition
Model
represents all superpixels within the time window of frames [t - , t]
represents all poses within the time window of frames [t - , t]
Online Localization and Prediction of Actions
Ø Given Pose bounding box and Superpixels
Ø Learn an Appearance model by clustering superpixels
Ø Each cluster is assigned a confidence based on the overlap area with the detection
Superpixel-based Foreground Likelihood Online Localization and Prediction of Actions
Ø Appearance model generates confidence for each superpixel
Ø Superpixel based Foreground Likelihood
Ø The Appearance Model is updated every 5th frame
Superpixel-based Foreground Likelihood Online Localization and Prediction of Actions
Ø Pose based Foreground Likelihood
Ø Refine Poses Ø Impose smoothness
constraints on joints: Ø Appearance Ø Location Ø Scale
Posed-based Foreground Likelihood Online Localization and Prediction of Actions
Posed-based Foreground Likelihood
• Minimize Cost Function: Raw Appearance Location Scale
Online Localization and Prediction of Actions
Posed-based Foreground Likelihood
• Appearance smoothness of joints: Appearance
Online Localization and Prediction of Actions
Posed-based Foreground Likelihood
• Location smoothness of joints:
Joint Location Difference
Location
Online Localization and Prediction of Actions
Posed-based Foreground Likelihood
• Scale smoothness of joints:
j'min
j'max
Scale
Online Localization and Prediction of Actions
Posed Refinement B
ef or
e A
ft er
Kicking Walking
Online Localization and Prediction of Actions
Ø Construct a Graph Ø Nodes = Superpixels Ø Edges = Spatio-Temporally
connected
Spatio-Temporal Graph Online Localization and Prediction of Actions
Ø Conditional Random Field Ø Unary Potential:
Ø Superpixel and Pose based Foreground Likelihood
Ø Binary Potential: Ø Color similarity Ø Flow and Motion Boundary Ø Edges
Conditional Random Field Online Localization and Prediction of Actions
Ø SVM Training Ø Divide videos into 1 second
temporal segments Ø Train SVM for each segment
Action Prediction using SVM
1 sec
2 sec
3 sec
1 sec
2 sec
3 sec
Online Localization and Prediction of Actions
Ø Action Prediction in Testing Video Ø Dynamic Programming
Ø Accumulate confidence of matching sequence
Ø Features Ø Improved Dense Trajectory
Features Ø Pose Features
Ø Normalized joint positions Ø Relative position of normalized joints Ø Orientation of vector connecting joints Ø Inner angle of vectors connecting two
joints
Action Prediction using SVM Online Localization and Prediction of Actions
Summary of Proposed Framework
(a) Input Stream of Video Frames
(b) Superpixel Extraction and Pose
Estimation
(c) Learn Superpixel based Appearance
Model
(d) Superpixel based Foreground Likelihood
(e) Pose Refinement
(f ) Segment Action with CRF + Action
Prediction using SVM
Online Localization and Prediction of Actions
Experimental Setup
• Features: • Superpixels: SLIC • Pose Estimation: Yang and Ramanan • Improved dense trajectories (iDTF) • Pose Features (joint locations, trajectories and orientations)
• Evaluation Metrics: • Receiver Operator Characteristic (ROC) curves • Area Under the Curve (AUC) • * Prediction Accuracy vs. Video Observation Percentage • * AUC vs. Video Observation Percentage
• Baseline • Exhaustively generate bounding boxes for each frame and connect them over
time using appearance similarity
Online Localization and Prediction of Actions
Experiment Datasets
• JHMDB • 21 Actions • 928 Videos
• UCF Sports • 10 actions • 150 videos
Online Localization and Prediction of Actions
Action Prediction + Detection
ØTesting Video Example 1: Kick Ball Action
Sorted Prediction Scores
Online Localization and Prediction of Actions
Ground Truth ( )
Action Localization Segment ( )
Action Localization Results Online Localization and Prediction of Actions
JHMDB
UCF Sports
0.2 0.4 0.6 0.8 1 0
0.1
0.2
0.3
0.4
0.5
0.6
0.7
0.8
Video Observation Percentage
A cc
ur ac
y
ØWe analyze how the prediction accuracy varies across time
0.2 0.4 0.6 0.8 1 0
0.1
0.2
0.3
0.4
0.5
0.6
0.7
0.8
Video Observation Percentage
A cc
ur ac
y
JHMDB
0.2 0.4 0.6 0.8 1 0
0.1
0.2
0.3
0.4
0.5
0.6
0.7
0.8
Video Observation Percentage
A cc
ur ac
y
JHMDB UCF Sports
Action Prediction with Time
Online Localization and Prediction of Actions
Action Prediction with Time
Ø We analyze the prediction accuracy for each action class of JHMDB Ø Actions are sorted by their prediction score, using AUC: - Ø Pullup Ø Golf Ø Pour Ø Push Ø Brush Hair Ø Climb Stairs Ø Clap Ø Shoot Bow Ø Sit Ø Run Ø Walk Ø Shoot Gun Ø Stand Ø Swing Baseball Ø Shoot Ball Ø Pick Ø Wave Ø Kick Ball Ø Catch Ø Throw Ø Jump
Easiest
Challenging
0.2 0.4 0.6 0.8 1 0
0.2
0.4
0.6
0.8
1
Video Observation Percentage
A cc
ur ac
y
0.2 0.4 0.6 0.8 1 0
0.2
0.4
0.6
0.8
1
Video Observation Percentage
A cc
ur ac
y
Brush Hair
0.2 0.4 0.6 0.8 1 0
0.2
0.4
0.6
0.8
1
Video Observation Percentage
A cc
ur ac
y
Brush Hair Catch
0.2 0.4 0.6 0.8 1 0
0.2
0.4
0.6
0.8
1
Video Observation Percentage
A cc
ur ac
y
Brush Hair Catch Clap
0.2 0.4 0.6 0.8 1 0
0.2
0.4
0.6
0.8
1
Video Observation Percentage
A cc
ur ac
y
Brush Hair Catch Clap Climb Stairs Golf Jump Kick Ball Pick Pour Pullup Push
Run Shoot Ball Shoot Bow Shoot Gun Sit Stand Swing Baseball Throw Walk Wave Mean Accuracy
Online Localization and Prediction of Actions
Action Localization with Time
ØAnalyze localization performance across time
JHMDB UCF-Sports
0.2
0.2 0.4 0.6 0.8 1 0
0.1
0.3
0.4
0.5
0.6
Video Observation Percentage
A U
C 0.2 0.4 0.6 0.8 1
0
0.1
0.2
0.3
0.4
0.5
0.6
Video Observation Percentage
A U
C
0.2
0.2 0.4 0.6 0.8 1 0
0.1
0.3
0.4
0.5
0.6
Video Observation Percentage
A U
C
10% 20%
30% 40%
50% 60%
0.2 0.4 0.6 0.8 1 0
0.1
0.2
0.3
0.4
0.5
0.6
Video Observation Percentage
A U
C
10% 20%
30% 40%
50% 60%
0.2
0.2 0.4 0.6 0.8 1 0
0.1
0.3
0.4
0.5
0.6
Video Observation Percentage
A U
C
10% 20%
30% 40%
50% 60%
0.2 0.4 0.6 0.8 1 0
0.1
0.2
0.3
0.4
0.5
0.6
Video Observation Percentage
A U
C
10% 20%
30% 40%
50% 60%
0.2
0.2 0.4 0.6 0.8 1 0
0.1
0.3
0.4
0.5
0.6
Video Observation Percentage
A U
C
10% 20%
30% 40%
50% 60%
0.2 0.4 0.6 0.8 1 0
0.1
0.2
0.3
0.4
0.5
0.6
Video Observation Percentage
A U
C
10% 20%
30% 40%
50% 60%
0.2
0.2 0.4 0.6 0.8 1 0
0.1
0.3
0.4
0.5
0.6
Video Observation Percentage
A U
C
10% 20%
30% 40%
50% 60%
0.2 0.4 0.6 0.8 1 0
0.1
0.2
0.3
0.4
0.5
0.6
Video Observation Percentage
A U
C
10% 20%
30% 40%
50% 60%
0.2
0.2 0.4 0.6 0.8 1 0
0.1
0.3
0.4
0.5
0.6
Video Observation Percentage
A U
C
10% 20%
30% 40%
50% 60%
0.2 0.4 0.6 0.8 1 0
0.1
0.2
0.3
0.4
0.5
0.6
Video Observation Percentage
A U
C
10% 20%
30% 40%
50% 60%
0.2
0.2 0.4 0.6 0.8 1 0
0.1
0.3
0.4
0.5
0.6
Video Observation Percentage
A U
C
10% 20%
30% 40%
50% 60%
0.2 0.4 0.6 0.8 1 0
0.1
0.2
0.3
0.4
0.5
0.6
Video Observation Percentage
A U
C
10% 20%
30% 40%
50% 60%
Online Localization and Prediction of Actions
Action Localization with Off line Methods
• Experiments on UCF Sports
0.1 0.2 0.3 0.4 0.5 0.6 0
0.1
0.2
0.3
0.4
0.5
0.6
Overlap Threshold
A U
C
0 0.1 0.2 0.3 0.4 0.5 0.6 0
0.2
0.4
0.6
0.8
1
False Positive Rate
T ru
e P
os it
iv e
R at
e
0.1 0.2 0.3 0.4 0.5 0.6 0
0.1
0.2
0.3
0.4
0.5
0.6
Overlap Threshold
A U
C
Lan et al. Tian et al. Wang et al. Gemert et al. Jain et al. (CVPR14) Jain et al. (CVPR15) Gkioxari and Malik
0 0.1 0.2 0.3 0.4 0.5 0.6 0
0.2
0.4
0.6
0.8
1
False Positive Rate
T ru
e P
o si
ti ve
R at
e
Lan et al. Tian et al. Jain et al. Wang et al. Gkioxari and Malik
0.1 0.2 0.3 0.4 0.5 0.6 0
0.1
0.2
0.3
0.4
0.5
0.6
Overlap Threshold
A U
C
Lan et al. Tian et al. Wang et al. Gemert et al. Jain et al. (CVPR14) Jain et al. (CVPR15) Gkioxari and Malik Baseline (Online)
0 0.1 0.2 0.3 0.4 0.5 0.6 0
0.2
0.4
0.6
0.8
1
False Positive Rate
T ru
e P
o si
ti ve
R at
e
Lan et al. Tian et al. Jain et al. Wang et al. Gkioxari and Malik Baseline (Online)
0.1 0.2 0.3 0.4 0.5 0.6 0
0.1
0.2
0.3
0.4
0.5
0.6
Overlap Threshold
A U
C
Lan et al. Tian et al. Wang et al. Gemert et al. Jain et al. (CVPR14) Jain et al. (CVPR15) Gkioxari and Malik Baseline (Online) Proposed
0 0.1 0.2 0.3 0.4 0.5 0.6 0
0.2
0.4
0.6
0.8
1
False Positive Rate
T ru
e P
o si
ti ve
R at
e
Lan et al. Tian et al. Jain et al. Wang et al. Gkioxari and Malik Baseline (Online) Proposed
Online Localization and Prediction of Actions
Summary – Online Localization
ØNew problem of Online Action Localization ØSimultaneous prediction and detection of actions
ØHigh-level poses and mid-level superpixels to compute foreground likelihood ØInfer the action location through Conditional Random Field ØDetailed analysis of prediction and localization across time
Online Localization and Prediction of Actions
Limitations
• Previous supervised approaches require: • Manual video class labeling • Manual bounding box annotations per frame
• Such efforts can be time consuming and expensive • Impractical for large number of classes and videos • Can lead to unwanted biases and errors
Online Localization and Prediction of Actions
Outline
I. Online Localization: Online Localization and Prediction of Actions and Interactions
II. Unsupervised Localization: Unsupervised Action Discovery and Localization in Videos
Unsupervised Action Discovery and Localization in Videos
Unsupervised Action Discovery and Localization in Videos
Soomro, Khurram and Mubarak Shah, ”Unsupervised Action Discovery and Localization in Videos”, Proceedings of the IEEE Conference on International Conference in Computer Vision, 2017.
Comparison: Supervised vs. Unsupervised
1. Video Action Class Labels
2. Spatio-Temporal Bounding box annotations
1. No Video Labels
2. No Annotations
Supervised Action Localization Unsupervised Action Localization
Problem Definition
We present a novel approach to solve three problems: 1. Action Discovery or Clustering in Training Videos 2. Action Annotation in Training Videos 3. Action Localization in Testing Videos
Unsupervised Action Discovery and Localization in Videos
Proposed Approach
Input
• Multiple Video Action Classes • No Video Labels • No Annotations
Output
1. Video Action Clusters 2. Action Annotations 3. Action Localizations
Unsupervised Action Discovery and Localization in Videos
Unsupervised Action Discovery and Localization in Videos
Action Discovery through Discriminative Clustering
• 0-1 Knapsack problem: Given a set of items, each with a weight and a value, determine the subset of items to include in a collection, so that the total weight is less than a given limit and total value is as high as possible.
Spatio-Temporal Annotation using Knapsack Unsupervised Action Discovery and Localization in Videos
Spatio-Temporal Annotation using Knapsack Unsupervised Action Discovery and Localization in Videos
Spatio-Temporal Annotation using Knapsack Unsupervised Action Discovery and Localization in Videos
• Knapsack Value:
Humanness
Saliency Motion Boundary
Spatio-Temporal Annotation using Knapsack Unsupervised Action Discovery and Localization in Videos
• Knapsack Weight:
1 2 3 4 5 6 7 8
-1 0 1 0 0 0 0 0 0 0!"#2
3
5
6
1
4
7
8
$ %
%$
0 0 0 0 0 -1 -1 1 0 0!"&
2
3
5
6
1
4
7
8
$ %
Video (()),SV /"0 , SV features 8"0
(b) Compute Knapsack value and weight
(a) Segment Video into Supervoxels (SVs)
(c) C9:;<=>?< @=ABC using SVs
(d) Temporal Constraints for Knapsack
(e) Select SVs using Knapsack Optimization
(g) Action Annotation
D)(E),F))Value G"0 ,Weight L"0 M) Knapsack B" N
Spatio-Temporal Annotation using Knapsack Unsupervised Action Discovery and Localization in Videos
• Directed Acyclic Graph:
• Temporal Constraints:
1 2 3 4 5 6 7 8
-1 0 1 0 0 0 0 0 0 0!"#2
3
5
6
1
4
7
8
$ %
%$
0 0 0 0 0 -1 -1 1 0 0!"&
2
3
5
6
1
4
7
8
$ %
Video (()),SV /"0 , SV features 8"0
(b) Compute Knapsack value and weight
(a) Segment Video into Supervoxels (SVs)
(c) C9:;<=>?< @=ABC using SVs
(d) Temporal Constraints for Knapsack
(e) Select SVs using Knapsack Optimization
(g) Action Annotation
D)(E),F))Value G"0 ,Weight L"0 M) Knapsack B" N
1 2 3 4 5 6 7 8
-1 0 1 0 0 0 0 0 0 0!"#2
3
5
6
1
4
7
8
$ %
%$
0 0 0 0 0 -1 -1 1 0 0!"&
2
3
5
6
1
4
7
8
$ %
Video (()),SV /"0 , SV features 8"0
(b) Compute Knapsack value and weight
(a) Segment Video into Supervoxels (SVs)
(c) C9:;<=>?< @=ABC using SVs
(d) Temporal Constraints for Knapsack
(e) Select SVs using Knapsack Optimization
(g) Action Annotation
D)(E),F))Value G"0 ,Weight L"0 M) Knapsack B" N
Spatio-Temporal Annotation using Knapsack Unsupervised Action Discovery and Localization in Videos
• Knapsack Formulation • Binary Integer Linear Programming (BILP)
1 2 3 4 5 6 7 8
-1 0 1 0 0 0 0 0 0 0!"#2
3
5
6
1
4
7
8
$ %
%$
0 0 0 0 0 -1 -1 1 0 0!"&
2
3
5
6
1
4
7
8
$ %
Video (()),SV /"0 , SV features 8"0
(b) Compute Knapsack value and weight
(a) Segment Video into Supervoxels (SVs)
(c) C9:;<=>?< @=ABC using SVs
(d) Temporal Constraints for Knapsack
(e) Select SVs using Knapsack Optimization
(g) Action Annotation
D)(E),F))Value G"0 ,Weight L"0 M) Knapsack B" N
1 2 3 4 5 6 7 8
-1 0 1 0 0 0 0 0 0 0!"#2
3
5
6
1
4
7
8
$ %
%$
0 0 0 0 0 -1 -1 1 0 0!"&
2
3
5
6
1
4
7
8
$ %
Video (()),SV /"0 , SV features 8"0
(b) Compute Knapsack value and weight
(a) Segment Video into Supervoxels (SVs)
(c) C9:;<=>?< @=ABC using SVs
(d) Temporal Constraints for Knapsack
(e) Select SVs using Knapsack Optimization
(g) Action Annotation
D)(E),F))Value G"0 ,Weight L"0 M) Knapsack B" N
1 2 3 4 5 6 7 8
-1 0 1 0 0 0 0 0 0 0!"#2
3
5
6
1
4
7
8
$ %
%$
0 0 0 0 0 -1 -1 1 0 0!"&
2
3
5
6
1
4
7
8
$ %
Video (()),SV /"0 , SV features 8"0
(b) Compute Knapsack value and weight
(a) Segment Video into Supervoxels (SVs)
(c) C9:;<=>?< @=ABC using SVs
(d) Temporal Constraints for Knapsack
(e) Select SVs using Knapsack Optimization
(g) Action Annotation
D)(E),F))Value G"0 ,Weight L"0 M) Knapsack B" N
Experiment Datasets
Action Discovery Experimental Results: 1. UCF Sports 2. JHMDB 3. Sub-JHMDB 4. THUMOS13 5. UCF101
Unsupervised Action Discovery and Localization in Videos
Experiment Datasets
Action Localization Experimental Results: 1. UCF Sports 2. JHMDB 3. Sub-JHMDB 4. THUMOS13
Unsupervised Action Discovery and Localization in Videos
Experimental Setup
• Features: • C3D for Action Discovery • iDTF (improved Dense Trajectory Features) for Action Localization
• Evaluation Metrics: • Action Discovery
• Percentage Clustering Accuracy • Action Localization
• Receiver Operator Characteristic (ROC) curves • Area Under the Curve (AUC)
Unsupervised Action Discovery and Localization in Videos
Unsupervised Action Discovery Results Unsupervised Action Discovery and Localization in Videos
Unsupervised Action Annotation Results Unsupervised Action Discovery and Localization in Videos
Weakly Supervised Action Localization:
Ground Truth ( )Action Localization Segment ( )
Action Localization Results Unsupervised Action Discovery and Localization in Videos
UCF Sports THUMOS 13
Ground Truth ( )Action Localization Segment ( )
Action Localization Results Unsupervised Action Discovery and Localization in Videos
JHMDB Sub-JHMDB
Unsupervised Action Localization Results Unsupervised Action Discovery and Localization in Videos
UCF Sports THUMOS 13
Unsupervised Action Localization Results Unsupervised Action Discovery and Localization in Videos
JHMDB Sub-JHMDB
Summary – Unsupervised Localization
üNew problem of Unsupervised Action Localization üAutomatic discovery of action class labels üNovel Knapsack approach to annotate actions in training videos
Unsupervised Action Discovery and Localization in Videos
Thank You!