Critique about (Action Localization in Videos )

profilevicky333
Talk_FIT_Khurram_Soomro.compressed1.pdf

Action Localization in Videos

Khurram Soomro Center for Research in Computer Vision (CRCV)

University of Central Florida (UCF)

Action Recognition

Diving Lifting

Golf

Swing Bench Walking

Action Localization

1. Action Recognition 2. Action Detection

a. Trimmed Videos i. Spatial

b. Untrimmed Videos i. Temporal ii. Spatio-Temporal

Diving

Lifting

Swing Bench

• Cluttered Background

• Multiple Actors/Actions

• Untrimmed Videos

Basketball Dunk

Salsa Spin

Hand Waving/Clapping/Boxing

Challenges: Action Localization

Applications of Action Localization

• Video Search • Action Retrieval • Multimedia Event Recounting • Video Understanding

Outline

I. Online Localization: Online Localization and Prediction of Actions

II. Unsupervised Localization: Unsupervised Action Discovery and Localization in Videos

Outline

I. Online Localization: Online Localization and Prediction of Actions

II. Unsupervised Localization: Unsupervised Action Discovery and Localization in Videos

Online Localization and Prediction of Actions

Online Localization and Prediction of Actions

Soomro, Khurram, Haroon Idrees, and Mubarak Shah, "Predicting the Where and What of actors and actions through Online Action Localization”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016.

Soomro, Khurram, Haroon Idrees, and Mubarak Shah, "Online Localization and Prediction of Actions and Interactions", IEEE PAMI, 2017 (Under Review).

Comparison: Off line vs. Online

1. Whole video available

2. Entire motion of action 3. Action Recognition

4. Action Detection using all frames

1. Partially observed video

2. Limited motion information 3. Action Prediction

4. Action Detection done frame-by-frame

Off line Action Localization Online Action Localization

Motivation

• Existing approaches localizes after an activity has occurred • Cannot anticipate actions • No timely localization of actions

Action Localization in Videos through Context Walk

Problem Definition

SORTED PREDICTION SCORES

Online Localization and Prediction of Actions

+ Recognition= Action DetectionOff line Action Localization

K IC

K IN

G + Prediction= Action DetectionOnline Action Localization

Problem Definition Online Localization and Prediction of Actions

Idea: Foreground/Background Distinction

✓Human Detection ✓Segmentation

✓Object ✓Motion ✓Action Proposals

✓Pose Estimation (Detecting Body Joints)

Online Localization and Prediction of Actions

Idea: Foreground/Background Distinction

✓Human Detection ✓Segmentation

✓Object ✓Motion ✓Action Proposals

✓Pose Estimation (Detecting Body Joints)

Online Localization and Prediction of Actions

Ø Given an input stream of video frames

Ø Goal: - Ø Predict action in current frame Ø Detect action in current frame

Ø Process in a batch of frames

Proposed Framework Online Localization and Prediction of Actions

Ø Extract Superpixels in each frame

Proposed Framework Online Localization and Prediction of Actions

Ø Pose Estimation

Proposed Framework Online Localization and Prediction of Actions

Framework: Online Action Localization

• Goal:

• Bayes Rule:

• State Transition model:

Pose-based Foreground Likelihood

= Location representing bounding box at time t

Superpixel-based Foreground Likelihood

State- Transition

Model

represents all superpixels within the time window of frames [t - , t]

represents all poses within the time window of frames [t - , t]

Online Localization and Prediction of Actions

Ø Given Pose bounding box and Superpixels

Ø Learn an Appearance model by clustering superpixels

Ø Each cluster is assigned a confidence based on the overlap area with the detection

Superpixel-based Foreground Likelihood Online Localization and Prediction of Actions

Ø Appearance model generates confidence for each superpixel

Ø Superpixel based Foreground Likelihood

Ø The Appearance Model is updated every 5th frame

Superpixel-based Foreground Likelihood Online Localization and Prediction of Actions

Ø Pose based Foreground Likelihood

Ø Refine Poses Ø Impose smoothness

constraints on joints: Ø Appearance Ø Location Ø Scale

Posed-based Foreground Likelihood Online Localization and Prediction of Actions

Posed-based Foreground Likelihood

• Minimize Cost Function: Raw Appearance Location Scale

Online Localization and Prediction of Actions

Posed-based Foreground Likelihood

• Appearance smoothness of joints: Appearance

Online Localization and Prediction of Actions

Posed-based Foreground Likelihood

• Location smoothness of joints:

Joint Location Difference

Location

Online Localization and Prediction of Actions

Posed-based Foreground Likelihood

• Scale smoothness of joints:

j'min

j'max

Scale

Online Localization and Prediction of Actions

Posed Refinement B

ef or

e A

ft er

Kicking Walking

Online Localization and Prediction of Actions

Ø Construct a Graph Ø Nodes = Superpixels Ø Edges = Spatio-Temporally

connected

Spatio-Temporal Graph Online Localization and Prediction of Actions

Ø Conditional Random Field Ø Unary Potential:

Ø Superpixel and Pose based Foreground Likelihood

Ø Binary Potential: Ø Color similarity Ø Flow and Motion Boundary Ø Edges

Conditional Random Field Online Localization and Prediction of Actions

Ø SVM Training Ø Divide videos into 1 second

temporal segments Ø Train SVM for each segment

Action Prediction using SVM

1 sec

2 sec

3 sec

1 sec

2 sec

3 sec

Online Localization and Prediction of Actions

Ø Action Prediction in Testing Video Ø Dynamic Programming

Ø Accumulate confidence of matching sequence

Ø Features Ø Improved Dense Trajectory

Features Ø Pose Features

Ø Normalized joint positions Ø Relative position of normalized joints Ø Orientation of vector connecting joints Ø Inner angle of vectors connecting two

joints

Action Prediction using SVM Online Localization and Prediction of Actions

Summary of Proposed Framework

(a) Input Stream of Video Frames

(b) Superpixel Extraction and Pose

Estimation

(c) Learn Superpixel based Appearance

Model

(d) Superpixel based Foreground Likelihood

(e) Pose Refinement

(f ) Segment Action with CRF + Action

Prediction using SVM

Online Localization and Prediction of Actions

Experimental Setup

• Features: • Superpixels: SLIC • Pose Estimation: Yang and Ramanan • Improved dense trajectories (iDTF) • Pose Features (joint locations, trajectories and orientations)

• Evaluation Metrics: • Receiver Operator Characteristic (ROC) curves • Area Under the Curve (AUC) • * Prediction Accuracy vs. Video Observation Percentage • * AUC vs. Video Observation Percentage

• Baseline • Exhaustively generate bounding boxes for each frame and connect them over

time using appearance similarity

Online Localization and Prediction of Actions

Experiment Datasets

• JHMDB • 21 Actions • 928 Videos

• UCF Sports • 10 actions • 150 videos

Online Localization and Prediction of Actions

Action Prediction + Detection

ØTesting Video Example 1: Kick Ball Action

Sorted Prediction Scores

Online Localization and Prediction of Actions

Ground Truth ( )

Action Localization Segment ( )

Action Localization Results Online Localization and Prediction of Actions

JHMDB

UCF Sports

0.2 0.4 0.6 0.8 1 0

0.1

0.2

0.3

0.4

0.5

0.6

0.7

0.8

Video Observation Percentage

A cc

ur ac

y

ØWe analyze how the prediction accuracy varies across time

0.2 0.4 0.6 0.8 1 0

0.1

0.2

0.3

0.4

0.5

0.6

0.7

0.8

Video Observation Percentage

A cc

ur ac

y

JHMDB

0.2 0.4 0.6 0.8 1 0

0.1

0.2

0.3

0.4

0.5

0.6

0.7

0.8

Video Observation Percentage

A cc

ur ac

y

JHMDB UCF Sports

Action Prediction with Time

Online Localization and Prediction of Actions

Action Prediction with Time

Ø We analyze the prediction accuracy for each action class of JHMDB Ø Actions are sorted by their prediction score, using AUC: - Ø Pullup Ø Golf Ø Pour Ø Push Ø Brush Hair Ø Climb Stairs Ø Clap Ø Shoot Bow Ø Sit Ø Run Ø Walk Ø Shoot Gun Ø Stand Ø Swing Baseball Ø Shoot Ball Ø Pick Ø Wave Ø Kick Ball Ø Catch Ø Throw Ø Jump

Easiest

Challenging

0.2 0.4 0.6 0.8 1 0

0.2

0.4

0.6

0.8

1

Video Observation Percentage

A cc

ur ac

y

0.2 0.4 0.6 0.8 1 0

0.2

0.4

0.6

0.8

1

Video Observation Percentage

A cc

ur ac

y

Brush Hair

0.2 0.4 0.6 0.8 1 0

0.2

0.4

0.6

0.8

1

Video Observation Percentage

A cc

ur ac

y

Brush Hair Catch

0.2 0.4 0.6 0.8 1 0

0.2

0.4

0.6

0.8

1

Video Observation Percentage

A cc

ur ac

y

Brush Hair Catch Clap

0.2 0.4 0.6 0.8 1 0

0.2

0.4

0.6

0.8

1

Video Observation Percentage

A cc

ur ac

y

Brush Hair Catch Clap Climb Stairs Golf Jump Kick Ball Pick Pour Pullup Push

Run Shoot Ball Shoot Bow Shoot Gun Sit Stand Swing Baseball Throw Walk Wave Mean Accuracy

Online Localization and Prediction of Actions

Action Localization with Time

ØAnalyze localization performance across time

JHMDB UCF-Sports

0.2

0.2 0.4 0.6 0.8 1 0

0.1

0.3

0.4

0.5

0.6

Video Observation Percentage

A U

C 0.2 0.4 0.6 0.8 1

0

0.1

0.2

0.3

0.4

0.5

0.6

Video Observation Percentage

A U

C

0.2

0.2 0.4 0.6 0.8 1 0

0.1

0.3

0.4

0.5

0.6

Video Observation Percentage

A U

C

10% 20%

30% 40%

50% 60%

0.2 0.4 0.6 0.8 1 0

0.1

0.2

0.3

0.4

0.5

0.6

Video Observation Percentage

A U

C

10% 20%

30% 40%

50% 60%

0.2

0.2 0.4 0.6 0.8 1 0

0.1

0.3

0.4

0.5

0.6

Video Observation Percentage

A U

C

10% 20%

30% 40%

50% 60%

0.2 0.4 0.6 0.8 1 0

0.1

0.2

0.3

0.4

0.5

0.6

Video Observation Percentage

A U

C

10% 20%

30% 40%

50% 60%

0.2

0.2 0.4 0.6 0.8 1 0

0.1

0.3

0.4

0.5

0.6

Video Observation Percentage

A U

C

10% 20%

30% 40%

50% 60%

0.2 0.4 0.6 0.8 1 0

0.1

0.2

0.3

0.4

0.5

0.6

Video Observation Percentage

A U

C

10% 20%

30% 40%

50% 60%

0.2

0.2 0.4 0.6 0.8 1 0

0.1

0.3

0.4

0.5

0.6

Video Observation Percentage

A U

C

10% 20%

30% 40%

50% 60%

0.2 0.4 0.6 0.8 1 0

0.1

0.2

0.3

0.4

0.5

0.6

Video Observation Percentage

A U

C

10% 20%

30% 40%

50% 60%

0.2

0.2 0.4 0.6 0.8 1 0

0.1

0.3

0.4

0.5

0.6

Video Observation Percentage

A U

C

10% 20%

30% 40%

50% 60%

0.2 0.4 0.6 0.8 1 0

0.1

0.2

0.3

0.4

0.5

0.6

Video Observation Percentage

A U

C

10% 20%

30% 40%

50% 60%

0.2

0.2 0.4 0.6 0.8 1 0

0.1

0.3

0.4

0.5

0.6

Video Observation Percentage

A U

C

10% 20%

30% 40%

50% 60%

0.2 0.4 0.6 0.8 1 0

0.1

0.2

0.3

0.4

0.5

0.6

Video Observation Percentage

A U

C

10% 20%

30% 40%

50% 60%

Online Localization and Prediction of Actions

Action Localization with Off line Methods

• Experiments on UCF Sports

0.1 0.2 0.3 0.4 0.5 0.6 0

0.1

0.2

0.3

0.4

0.5

0.6

Overlap Threshold

A U

C

0 0.1 0.2 0.3 0.4 0.5 0.6 0

0.2

0.4

0.6

0.8

1

False Positive Rate

T ru

e P

os it

iv e

R at

e

0.1 0.2 0.3 0.4 0.5 0.6 0

0.1

0.2

0.3

0.4

0.5

0.6

Overlap Threshold

A U

C

Lan et al. Tian et al. Wang et al. Gemert et al. Jain et al. (CVPR14) Jain et al. (CVPR15) Gkioxari and Malik

0 0.1 0.2 0.3 0.4 0.5 0.6 0

0.2

0.4

0.6

0.8

1

False Positive Rate

T ru

e P

o si

ti ve

R at

e

Lan et al. Tian et al. Jain et al. Wang et al. Gkioxari and Malik

0.1 0.2 0.3 0.4 0.5 0.6 0

0.1

0.2

0.3

0.4

0.5

0.6

Overlap Threshold

A U

C

Lan et al. Tian et al. Wang et al. Gemert et al. Jain et al. (CVPR14) Jain et al. (CVPR15) Gkioxari and Malik Baseline (Online)

0 0.1 0.2 0.3 0.4 0.5 0.6 0

0.2

0.4

0.6

0.8

1

False Positive Rate

T ru

e P

o si

ti ve

R at

e

Lan et al. Tian et al. Jain et al. Wang et al. Gkioxari and Malik Baseline (Online)

0.1 0.2 0.3 0.4 0.5 0.6 0

0.1

0.2

0.3

0.4

0.5

0.6

Overlap Threshold

A U

C

Lan et al. Tian et al. Wang et al. Gemert et al. Jain et al. (CVPR14) Jain et al. (CVPR15) Gkioxari and Malik Baseline (Online) Proposed

0 0.1 0.2 0.3 0.4 0.5 0.6 0

0.2

0.4

0.6

0.8

1

False Positive Rate

T ru

e P

o si

ti ve

R at

e

Lan et al. Tian et al. Jain et al. Wang et al. Gkioxari and Malik Baseline (Online) Proposed

Online Localization and Prediction of Actions

Summary – Online Localization

ØNew problem of Online Action Localization ØSimultaneous prediction and detection of actions

ØHigh-level poses and mid-level superpixels to compute foreground likelihood ØInfer the action location through Conditional Random Field ØDetailed analysis of prediction and localization across time

Online Localization and Prediction of Actions

Limitations

• Previous supervised approaches require: • Manual video class labeling • Manual bounding box annotations per frame

• Such efforts can be time consuming and expensive • Impractical for large number of classes and videos • Can lead to unwanted biases and errors

Online Localization and Prediction of Actions

Outline

I. Online Localization: Online Localization and Prediction of Actions and Interactions

II. Unsupervised Localization: Unsupervised Action Discovery and Localization in Videos

Unsupervised Action Discovery and Localization in Videos

Unsupervised Action Discovery and Localization in Videos

Soomro, Khurram and Mubarak Shah, ”Unsupervised Action Discovery and Localization in Videos”, Proceedings of the IEEE Conference on International Conference in Computer Vision, 2017.

Comparison: Supervised vs. Unsupervised

1. Video Action Class Labels

2. Spatio-Temporal Bounding box annotations

1. No Video Labels

2. No Annotations

Supervised Action Localization Unsupervised Action Localization

Problem Definition

We present a novel approach to solve three problems: 1. Action Discovery or Clustering in Training Videos 2. Action Annotation in Training Videos 3. Action Localization in Testing Videos

Unsupervised Action Discovery and Localization in Videos

Proposed Approach

Input

• Multiple Video Action Classes • No Video Labels • No Annotations

Output

1. Video Action Clusters 2. Action Annotations 3. Action Localizations

Unsupervised Action Discovery and Localization in Videos

Unsupervised Action Discovery and Localization in Videos

Action Discovery through Discriminative Clustering

• 0-1 Knapsack problem: Given a set of items, each with a weight and a value, determine the subset of items to include in a collection, so that the total weight is less than a given limit and total value is as high as possible.

Spatio-Temporal Annotation using Knapsack Unsupervised Action Discovery and Localization in Videos

Spatio-Temporal Annotation using Knapsack Unsupervised Action Discovery and Localization in Videos

Spatio-Temporal Annotation using Knapsack Unsupervised Action Discovery and Localization in Videos

• Knapsack Value:

Humanness

Saliency Motion Boundary

Spatio-Temporal Annotation using Knapsack Unsupervised Action Discovery and Localization in Videos

• Knapsack Weight:

1 2 3 4 5 6 7 8

-1 0 1 0 0 0 0 0 0 0!"#2

3

5

6

1

4

7

8

$ %

%$

0 0 0 0 0 -1 -1 1 0 0!"&

2

3

5

6

1

4

7

8

$ %

Video (()),SV /"0 , SV features 8"0

(b) Compute Knapsack value and weight

(a) Segment Video into Supervoxels (SVs)

(c) C9:;<=>?< @=ABC using SVs

(d) Temporal Constraints for Knapsack

(e) Select SVs using Knapsack Optimization

(g) Action Annotation

D)(E),F))Value G"0 ,Weight L"0 M) Knapsack B" N

Spatio-Temporal Annotation using Knapsack Unsupervised Action Discovery and Localization in Videos

• Directed Acyclic Graph:

• Temporal Constraints:

1 2 3 4 5 6 7 8

-1 0 1 0 0 0 0 0 0 0!"#2

3

5

6

1

4

7

8

$ %

%$

0 0 0 0 0 -1 -1 1 0 0!"&

2

3

5

6

1

4

7

8

$ %

Video (()),SV /"0 , SV features 8"0

(b) Compute Knapsack value and weight

(a) Segment Video into Supervoxels (SVs)

(c) C9:;<=>?< @=ABC using SVs

(d) Temporal Constraints for Knapsack

(e) Select SVs using Knapsack Optimization

(g) Action Annotation

D)(E),F))Value G"0 ,Weight L"0 M) Knapsack B" N

1 2 3 4 5 6 7 8

-1 0 1 0 0 0 0 0 0 0!"#2

3

5

6

1

4

7

8

$ %

%$

0 0 0 0 0 -1 -1 1 0 0!"&

2

3

5

6

1

4

7

8

$ %

Video (()),SV /"0 , SV features 8"0

(b) Compute Knapsack value and weight

(a) Segment Video into Supervoxels (SVs)

(c) C9:;<=>?< @=ABC using SVs

(d) Temporal Constraints for Knapsack

(e) Select SVs using Knapsack Optimization

(g) Action Annotation

D)(E),F))Value G"0 ,Weight L"0 M) Knapsack B" N

Spatio-Temporal Annotation using Knapsack Unsupervised Action Discovery and Localization in Videos

• Knapsack Formulation • Binary Integer Linear Programming (BILP)

1 2 3 4 5 6 7 8

-1 0 1 0 0 0 0 0 0 0!"#2

3

5

6

1

4

7

8

$ %

%$

0 0 0 0 0 -1 -1 1 0 0!"&

2

3

5

6

1

4

7

8

$ %

Video (()),SV /"0 , SV features 8"0

(b) Compute Knapsack value and weight

(a) Segment Video into Supervoxels (SVs)

(c) C9:;<=>?< @=ABC using SVs

(d) Temporal Constraints for Knapsack

(e) Select SVs using Knapsack Optimization

(g) Action Annotation

D)(E),F))Value G"0 ,Weight L"0 M) Knapsack B" N

1 2 3 4 5 6 7 8

-1 0 1 0 0 0 0 0 0 0!"#2

3

5

6

1

4

7

8

$ %

%$

0 0 0 0 0 -1 -1 1 0 0!"&

2

3

5

6

1

4

7

8

$ %

Video (()),SV /"0 , SV features 8"0

(b) Compute Knapsack value and weight

(a) Segment Video into Supervoxels (SVs)

(c) C9:;<=>?< @=ABC using SVs

(d) Temporal Constraints for Knapsack

(e) Select SVs using Knapsack Optimization

(g) Action Annotation

D)(E),F))Value G"0 ,Weight L"0 M) Knapsack B" N

1 2 3 4 5 6 7 8

-1 0 1 0 0 0 0 0 0 0!"#2

3

5

6

1

4

7

8

$ %

%$

0 0 0 0 0 -1 -1 1 0 0!"&

2

3

5

6

1

4

7

8

$ %

Video (()),SV /"0 , SV features 8"0

(b) Compute Knapsack value and weight

(a) Segment Video into Supervoxels (SVs)

(c) C9:;<=>?< @=ABC using SVs

(d) Temporal Constraints for Knapsack

(e) Select SVs using Knapsack Optimization

(g) Action Annotation

D)(E),F))Value G"0 ,Weight L"0 M) Knapsack B" N

Experiment Datasets

Action Discovery Experimental Results: 1. UCF Sports 2. JHMDB 3. Sub-JHMDB 4. THUMOS13 5. UCF101

Unsupervised Action Discovery and Localization in Videos

Experiment Datasets

Action Localization Experimental Results: 1. UCF Sports 2. JHMDB 3. Sub-JHMDB 4. THUMOS13

Unsupervised Action Discovery and Localization in Videos

Experimental Setup

• Features: • C3D for Action Discovery • iDTF (improved Dense Trajectory Features) for Action Localization

• Evaluation Metrics: • Action Discovery

• Percentage Clustering Accuracy • Action Localization

• Receiver Operator Characteristic (ROC) curves • Area Under the Curve (AUC)

Unsupervised Action Discovery and Localization in Videos

Unsupervised Action Discovery Results Unsupervised Action Discovery and Localization in Videos

Unsupervised Action Annotation Results Unsupervised Action Discovery and Localization in Videos

Weakly Supervised Action Localization:

Ground Truth ( )Action Localization Segment ( )

Action Localization Results Unsupervised Action Discovery and Localization in Videos

UCF Sports THUMOS 13

Ground Truth ( )Action Localization Segment ( )

Action Localization Results Unsupervised Action Discovery and Localization in Videos

JHMDB Sub-JHMDB

Unsupervised Action Localization Results Unsupervised Action Discovery and Localization in Videos

UCF Sports THUMOS 13

Unsupervised Action Localization Results Unsupervised Action Discovery and Localization in Videos

JHMDB Sub-JHMDB

Summary – Unsupervised Localization

üNew problem of Unsupervised Action Localization üAutomatic discovery of action class labels üNovel Knapsack approach to annotate actions in training videos

Unsupervised Action Discovery and Localization in Videos

Thank You!