cognitive psychology short essay
11/18/2019
1
Computational Object Recognition
Why is object recognition hard? • Hard to define what constitutes a particular object
• Edge ownership
• Occlusion
• Viewpoint invariance
• Fuzzy boundaries
11/18/2019
2
Object Recognition
• Determining what constitutes an object is a good start
• Once an object is defined, how do we recognize what it is?
• Object constancy
• Note that it’s not simply feedforward either – the identity of an object can help us classify it • If you are looking for a specific object (like your cell
phone), you might find and recognize it based on very minimal information
Visual Indeterminacy
• Before object recognition occurs, object is “indeterminant”
11/18/2019
3
11/18/2019
4
Viewpoint Invariance
• We are able to recognize objects regardless of what angle (viewpoint) we see the object at
• A good object recognizer needs to be viewpoint invariant
• There are limits to our abilities, but we are generally much better than might be expected given the level of success (well, failure) of classic computational object recognition algorithms
What are these objects?
11/18/2019
5
Recognition by Components
• Also called recognition by parts (Biederman)
• Segment an object into component “parts” • Also known as geons
• Sometimes called “geon theory”
• Determine position and relation between parts
• Match information with mental representation
11/18/2019
6
Recognition by Components
• Geons are viewpoint invariant • Big advantage to this class of theories
• Relation between geons holds true for all viewpoints
• As long as geons are visible
• Can be determined computationally
• Faster object naming when geon information is available
11/18/2019
7
Much easier when geons can be seen!
11/18/2019
8
Recognition by Views
• Different views of same object are stored in LTM • Exemplars
• Objects are then matched with exemplars
• Best match is the object
Canonical Object Views
• Most “obvious” viewpoint • Reflected in reaction time when naming picture
• What you would likely draw
• Mental image
• Recognition by Views • Canonical viewpoints are easily matched (you’ve seen
the object in that view the most)
• Recognition by Components • Canonical viewpoints have easily identifiable geons and
clearly show relation between them
These are also canonical though…
11/18/2019
9
RBC and RBV • Are they mutually exclusive?
• Some things might be better recognized using RBC • basic level objects (“a dog”)
• Can we have exemplars for abstract categories? (maybe)
• Some things might be better recognized using RBV • Specific instances of category (“my dog, Flapjack”)
• RBC can’t really recognize specific members
• Discriminating between two faces – geons are the same
Contextual Cues
• Something RBC and RBV don’t explicitly account for – effect of context on recognition
• What is the middle symbol in the picture? • A B C? 12 13 14? • Depends on top-down
contextual factors / pattern recognition
• Where do we include context into object recognition models? • Humans are very fast at
determining gist information about a scene – well before stable object recognition
What is this picture of?
11/18/2019
10
Context-constrained recognition • Ecological validity?
• Do lab experiments that involve object recognition in isolation from environment neglect constraining effects of environment on object possibilities?
• If gist comes before object recognition, then gist should be used to constrain recognition
• Also, consider than object recognition in the environment involves much more than just visual gist – higher level information about where you are, what you expect to see, etc.
Galleguillos & Belongie, 2010
11/18/2019
11
Contextual Cues
• Semantic context (probability) • Objects expected to be in a particular scene (given gist information)
• Spatial context (position) • Where objects are typically located • Location of objects constrains what they can (usually) be
• Scale context (size) • Relative size constrains what the objects are
• Basic idea – top-down knowledge can be utilized to simplify problem…but… • Hard to program computationally, as you need a computer with
sufficient top-down knowledge of the world • Some success though, especially with small-scale models where
specific contextual knowledge is “hard-coded” into the model
Galleguillos & Belongie, 2010
Connectionism and Object Recognition • Feature-based object recognition models, like RBC, lend
themselves to connectionist modeling
• Deep learning networks • Input layer (pixels), multiple hidden layers (features), output
layer (object identity) • Structure of hidden layer features NOT defined a priori –
system will “develop” it’s own representation of features at hidden layers through associative learning (more later)
• Associate input (picture) with output (identity)
• These models are quite successful and gaining a lot of traction in AI research
Comparing Humans and Deep Networks • Since deep learning networks are connectionist
networks, they might mirror how humans do object recognition
• Certain transformations are hard for both humans AND deep networks – overlap in processing similarity? • Rotation by depth – hardest
• Position – easiest
Kheradpisheh, Ghodrati, Ganjtabesh, & Masquelier, 2016
11/18/2019
12
Face Recognition
Face Perception
• We appear to be highly specialized to detect and process faces
• Even newborn infants show a preference for “face-like” stimuli
11/18/2019
13
Holistic Processing
• We process faces “holistically” • Which means we process it as
a whole object, NOT as a sum of its parts
• Only for upright faces (inversion effects)
• Galton, 1879
Hollow Mask Illusion
11/18/2019
14
Hollow Mask Illusion
• Objects are generally convex
• Faces are ALWAYS convex
• Faces processed holistically
• One step further – combine structure from motion with hollow mask illusion
Face Adaptation
• Adapting to a face can lead to a face aftereffect • Same process (basically)
as motion and color aftereffects…maybe
• Behold, Bushbama
11/18/2019
15
How do these faces differ?
11/18/2019
16
Face Space
• Multidimensional representation of all faces a person have ever seen • Exemplars (RBV) • Note that image is only 2
dimensions, actual face space is waaaaaaay more
• Dimensions include things like… • Distance between eyes • Size of nose • Width of mouth • Expanded/contracted features • Etc. etc. etc. etc.
Most “average” face
Highly distinctive face
Highly distinctive face
Face Space
• Face adaptation works because we compensate in an opposite direction across face space
• Adapt to one feature – next face looks “opposite” in those features
11/18/2019
17
Fusiform Face Area (FFA)
• Located in the right temporal lobe
• Selective for faces
• However, also active when car experts identify particular cars
• Left FFA highly active during reading
• Visual expertise in very similar stimuli? • Exemplars…? RBV?!?
11/18/2019
18
Greebles!
• Highly homogeneous artificial stimuli used to study object classification similar to facial recognition
• When people become “expert” Greeble recognizers….
• The FFA is active during Greeble perception!
Gauthier, 1997
Autism and Recognition • Another odd note – facial
recognition is impaired among people with ASD
• This impairment extends to detailed visual categorization as well (Greebles)
• Normal object recognition not impaired
• How is this related to avoidance of eye contact? Cause or symptom? • Clearly, prosopagnosia like deficits
would make social interaction more difficult
• Also likely related to difficulty appraising emotional expressions
Scherf, Behrmann, Minshew, & Luna, 2008
Prosopagnosia • If the FFA is better conceptualized as being utilized for
any visual expertise task… • Do people with prosopagnosia (inability to recognize
familiar faces) show deficits in other visual expertise tasks?
• Objects • Acquired prosopagnosics show impaired expertise-adjusted
object recognition (again, cars for car experts) • Congenital prosopagnosics don’t appear to have the same
deficits – why???
• Words • Mostly intact – note that VWFA is left hemisphere
corresponding location to FFA
• Also thought that prosopagnosia can be heterogenous – apperceptive and associative varieties, and differences between acquired and congenital. Also deficits can follow a continuum, not always binary (especially for congenital)
Corrow, Dalrymple, & Barton, 2016
11/18/2019
19
Bringing it all Back
• Face recognition is not necessarily an entirely special process – but it is a form of expert object recognition and does appear to be somewhat privileged
• RBC is insufficient to explain face recognition
• Affordances don’t seem to contribute to understanding face recognition?
• RBV, and specifically exemplars, seem to be required to perceive individual faces as unique
Embodied Object Recognition
Affordances • Every object in the world can be
acted on in some way
• Object affordances – potential actions you can take on an object • How your body can interact with an
object – a property determined by both the object and the body
• A mug affords holding and drinking (moving arm to mouth) • A pencil affords grasping and writing • A big red button affords pushing
• Object affordances, according to Gibson, are perceived directly • No additional processing needed –
direct perception Gibson, 1979 (and much more)
11/18/2019
20
Affordance Interpretation • Object perception and recognition also involves
knowing what actions are taken on that object
• Knowing an objects identity means knowing its affordances
• But does knowing affordances help recognition?
• Correlative link between identity of object and object affordance – are mirror neurons affordance sensitive?
Affordances
• Are affordances learned through interaction or based solely on innate bodily representation?
• Chemero – affordances are relations between bodily ability and environmental features
• Affordances are automatically acted upon without attentional control (unless you need to inhibit)
• Kirsh- goals make perception enactive – object recognition (categorization/use) depend on current goals – example of bricks in bricklayer vs. vandal
11/18/2019
21
Utilization behavior • Inability to inhibit acting on nearby
objects • Not conscious of action taken on
objects • Action aligns with object affordances
• Seen with damage or impairment to frontal lobes • EF disorders • Schizophrenia • Infants / children
• What does this tell us about affordances? • We can act on an objects’
affordances without being conscious of the action
• Action inhibition and inhibiting affordance based behavior requires consciousness?
Lhermitte, 1983
Anarchic Hand (Alien Hand)
• Loss of conscious control of a hand
• “Anarchic” hand will engage in utilization behavior on its own • One patient – anarchic
hand kept unbuttoning shirt as controlled hand was buttoning it
• Another patient – anarchic hand picked up fish bones off a friends plate and shoved them into her mouth
11/18/2019
22
https://www.youtube.com/watch?v=G3JceitLE2g
Embodied Object Recognition
• Do affordances cue object recognition? • Are they necessary for object recognition?
• Gibson – we directly perceive object affordances • If affordances are directly perceived, then this can form the
basis of object recognition without representation
• To know something is “a hammer” requires knowledge of experience using a hammer • Sensorimotor processes involved with hammer use are
inherent in object definition and thus object recognition/categorization
• Affordances play a role in object classification, especially for objects that are acted on
Associative Agnosia
• Damage in associative agnosia is quite specific to visual features of objects
• Object affordances still intact – and can be identified through action (remember the guy turning the lock?)
• Objects can be identified through haptic perception
• Auditory object recognition still possible
11/18/2019
23
Predictions • Visual presentation of manipulable objects activates
motor programs • Yes, but relationship is correlational – unclear if motor activity
is necessary for object recognition or just primed • Utilization behavior – without inhibition, motor program
activation manifests as actual action
• Motor activity is required for object naming • Double dissociation in naming ability between manipulable
and non-manipulable objects • Visual apraxia – deficit in action production of objects
• Mixed evidence, but some patients with apraxia have deficits in naming and describe manipulable objects
• Only SOME patients though - evidence against strong EC model – but unclear if this is due to plasticity following acquisition of apraxia
• Performing an unrelated action slows naming of manipulable objects…sometimes (more mixed evidence here)
Matheson, White, & McMullen, 2015
Predictions
• Existence of distinct neural assemblies integrating action and vision (cued by either action or vision) • Mirror neuron assemblies, but we discussed issues here
• Relationship between visual and motor association develops based on experience with objects • Children as young as 4 prefer to categorize based on function
and not visual similarity • Children are faster at matching manipulable thematic
relationships (screwdriver goes with screw) over non- manipulable (castle goes with knight)
• Implication – role of action experience in grounding conceptual representation is greater in early development (when direct interaction with world is more central to cognition?)
Matheson, White, & McMullen, 2015
Predictions
• Visual presentation of objects or actions primes tasks involving either vision or action • Priming with a correctly oriented
visual stimulus speeds RT for grasping subsequent bar when presented when comparted to priming improperly oriented stimulus
• Possibly also explained through cuing attention
• Association seen – but again, not necessarily a causal association
Matheson, White, & McMullen, 2015
11/18/2019
24
Evaluating EC and Object Recognition • Clearly a relationship between object recognition and
affordances, but… • Relationship does NOT appear to be causal and required (i.e.,
you don’t need to simulate action in order to perceive/recognize/name object)
• Embodied object recognition theories don’t have more/better evidence than computational theories
• Visual information alone appears sufficient to identify objects in most cases
• If anything, a weak form of EC seems more applicable here
• Still a lot of questions and mixed evidence – and of course the possibility that some redundant coding is present
• Like we mentioned with mirror neurons – perhaps it’s more about learning affordance -> identity association
One other note…
• Remember the dual visual system theory?
• Split between “where” and “what” / dorsal and ventral
• Is this a point against embodied object recognition? • Implies identity of an object is unrelated to its use (since
motor cortex is part of dorsal “where” stream, not ventral “what” stream)
• You’ll be revisiting this theory in your readings for this lecture!
Categorization
11/18/2019
25
Why Categorize?
• Organize a complicated, messy world
• Many regularities in the environment
• Generalization and discrimination
• Prediction – if an object is determined to be a member of a category, it can be treated like others in that category
Dark side of categorization
• Stereotyping • Judgments about
people determined by category membership
• Prejudice and bias
• Incorrect prediction • Could have disastrous
consequences!
Concepts vs. Categories
• Concept • Mental representation of a particular class • What is being represented
• Category • Set of entities grouped together • External to mental representation
• Which comes first? • Does your concept of house constrain what you
categorize as a “house”? • Or does what you categorize as a “house” determine
your concept of house?
Smith, 1989
11/18/2019
26
Categorizing Stimuli
• How do we classify objects?
• This is a harder problem than it might appear…
Necessary and Sufficient Conditions • Qualities that an item must have to be in a category as well
as discriminate from other category membership
• A cup is an object with a cylindrical body and an open top. The top of the cup is wider than the bottom, and it is taller than it is wide. If the object is similar, but shorter than it is wide, it is a bowl.
• Seems pretty good right?
11/18/2019
27
Functional Hypothesis
• Of course, we can continue to revise necessary and sufficient conditions, but they typically break down at some point.
• Maybe it’s more about the function of the object; a cup is used for holding and drinking liquid.
• Categories defined by their affordances
Causal History • Sometimes what matters most for categorization is
not similarity or function, but history • What makes a piece of art a “Picasso”? • What makes counterfeit money different from regular
money?
Functional vs. Historical • When is function more important
and when is history more important?
• Depends on point of categorization
• For man-made objects, tools, manipulable objects – probably function • Side note – as mentioned last time,
identifying/categorizing manipulable objects is correlated with left PMC activity (e.g., Gerlach, Law, & Paulson, 2002)
• For biological entities, or objects where history matters – probably history
• What about just visual similarity?
11/18/2019
28
Quick! Imagine a car!
• How many wheels does it have?
• What color is it?
• What style is it?
Prototype Theory
• Maybe it’s not strict rules, function, or history, but similarity, as in a family resemblance
• Every concept is typified by a certain prototypical instance of that concept
• The car you just pictured was prototypical (probably)
• Things that are more similar to the prototype are better members of that category
• “Fuzzy” category boundaries (as opposed to strict)
Rosch, 1971
11/18/2019
29
The Prototypical Bird
• Sparrow
• Crow
• Pelican
• Flamingo
• Penguin
• Certain birds feel more “bird-like”
11/18/2019
30
Typicality Ratings • Prototypical members can be determined through typicality
ratings
• Rate the following on a 1-7 scale for how much they fit into the category “fruit” (psychologically; these are all fruits) • Peach • Apple • Fig • Pomegranate • Blueberry • Olive • Strawberry • Raisin • Watermelon • Pumpkin • Pear • Lemon • Grape • Avocado
Posner & Keele, 1968 Demo
11/18/2019
31
Fuzzy Boundaries…good?
• One of main strengths of prototype theory is fuzzy boundaries – don’t need to abide by strict rules and category membership
• Problem – how fuzzy are the boundaries?
• How atypical can we move things from a prototype before it no longer matches?
Exemplar Theory
• Similar to prototype theory – compare new instance with representation
• Prototype theory – representation is “idealized” and “averaged” across all category members • Might not actually correspond to anything in reality! • Abstract representation generated by everything seen • A typical member of the category generated internally • Basically RBC
• Exemplar theory – representation corresponds to actual instance of category member • Exemplar of a category existed at one point and you saw it • “Best example(s)” of the category you have actually perceived • Collection of numerous specific members of the category • Basically RBV
11/18/2019
32
Category Boundaries
• Another way of thinking about categorization – place emphasis on boundaries between categories
• Caricatures of categories often quicker to be categorized than prototypical members • Caricatures are moved in direction away from possible
category boundary
Goldstone et al., 2012