Tearm paper

profileMesfer Alotaibi
lecture_5_chapter_5_1.ppt

Operant Reinforcement

Chapter 5

Introduction

While Pavlov was deciphering the psychic reflex, E.L. Thorndike was working on comprehending animal intelligence.

How do we study and measure animal intelligence?

The Law of Effect

B.F. Skinner coined operant learning based on Thorndike’s work.

Before Thorndike’s work, people believed animals had the ability to reason because of anecdotal evidence of unique events that were thought to generalize to all animals (e.g. a single cat finding its way home from miles away). Thorndike thought that this was nonsensical because single anecdotes certainly do not generalize. Instead of thinking that animals have the ability to reason, which was difficult to measure, Thorndike decided that animal intelligence only could be studied and measured using learning procedures. This involves giving an animal a task repeatedly and then measuring performance on the task. Thorndike, E. (1932). The Fundamentals of Learning. New York: Teachers College Press.

Thorndike’s famous experiment involved measuring learning in cats using the puzzle box in Figure 5-1. A hungry cat was placed in the box with food in plain view but out of reach. The door of the box could be opened by a simple act such as pulling on a wire loop. On the first trial, cats engaged in a number of ineffective acts, some of which seemed to result in frustration. However, with repeated trials, the animal learned to immediately escape. See the learning curve in Figure 5-2.

Law of Effect – There are two consequences to behavior that Thorndike called a satisfying state of affairs and an annoying state of affairs. This relationship between behavior and its consequences is known as the Law of Effect. The law states that the strength of behavior depends on the consequences for that behavior in the past. Prior to Thorndike, learning was thought to be a matter of reasoning. Thorndike shifted our attention to an organism’s external environment.

Skinner Box Procedure (Figure 5-3) – In the box there’s a food magazine that automatically drops food pellets into a tray. Near the food tray there is a lever that dispenses food pellets when pressed. A rat placed in the box is monitored to see at what point it discovers that pressing the lever results in dispensing a food pellet. Once the rat discovered the lever, lever pressing increased dramatically.

These procedures, whereby behavior is strengthened or weakened due to consequences, became known as operant learning because behavior is said to operate on the environment. Operant learning is not conditioning. It does not involve reflexes; it is more complicated. The organism operates on the environment to change it, and it is this change that either strengthens or weakens behavior.

Basic Procedures

Positive Reinforcement – Behavior followed by something the organism wants; therefore, it is strengthened or maintained.

B(behavior) --------------- SR (reinforcing stimulus)

Sit ------------------------ Bone

Negative Reinforcement – Behavior is strengthened by the removal or decrease in intensity of a negative stimulus.

B(behavior) --------------- SR (reinforcing stimulus)

Go through dog door ------- Escape rain

BOTH MAINTAIN / STRENGTHEN BEHAVIOR !!

Reinforcement provides a consequence for a behavior that maintains or increases its strength. A procedure must have three characteristics to qualify reinforcement:

1. The behavior ,must have a consequence,

2. The behavior must increase in strength, and

3. The increase in strength must be the result of the consequence.

There are two types of reinforcement: Positive and Negative reinforcement.

Positive Reinforcement – When a behavior is followed by the appearance of, of an increase in the intensity of, a stimulus, which increases the likelihood of the behavior in the future, the stimulus that appears is called a positive reinforcer, which is ordinarily something that an organism seeks out. Regarding positive reinforcement, the “something” the organism wants can be novel or it can be an increase in intensity of something already present. Positive reinforcement sometimes is called reward training. Remember, however, the same things are not rewarding for everyone!!

Negative Reinforcement – When a behavior is followed by the removal of, of a decrease in the intensity of, a stimulus, which increases the likelihood of the behavior in the future, the stimulus that appears is called a negative reinforcer, which is ordinarily something that an organism seeks to escape or avoid. It sometimes is called escape training.

BOTH MAINTAIN / STRENGTHEN BEHAVIOR !! The difference is that positive involves the appearance of a stimulus and negative involves the removal of a stimulus.

Get class to give examples of the B ----- SR relationship. Use human examples.

Discrete VS. Free Operant Procedures

Skinner and Thorndike studied operant procedures differently.

Discrete Trial Procedure (Thorndike) – Occurrence of the behavior ends the trial.

Free Operant Procedure (Skinner) – Behavior can be repeated an indefinite number of times.

The artificiality of laboratory studies is necessary for simplifying the problem to gain comprehension.

Discrete Trial Procedures (Thorndike) – Once the cat finds its way out of the puzzle box or once the rat finds its way out of the maze, the trial is over.

What we are measuring, with regard to learning, using this technique is time to complete the task, errors emitted, etc. (DV).

Free Operant Procedures (Skinner) – Measure learning usually through counting how many times the behavior occurs in a given time period, such as lever pressing. Because this can on all day and there is no beginning or end, we count the number of times the behavior occurs in a given time, instead.

In a clinical setting, discrete methods are used widely compared to free methods, although free methods are gaining popularity because they are considered a more natural approach. Discrete method example – Teaching a child with a lisp to repeat a word said by the teacher. Each time the word is repeated, there is a consequence, and the trial ends. Free method example – Teaching a child with a lisp by carrying a conversation while playing a game and providing consequences when the child says the target word, but continuing the conversation upon hearing the target word.

Lab studies are artificial, but they are necessary because problems must be simplified to identify functional relationships between IVs and DVs. Once relationships are identified, researchers can predict and control various phenomena in future studies.

Classical Vs. Operant learning

Contingency: US on CS vs. Stimulus on Behavior

Involuntary Reflexive Behavior vs. Voluntary Behavior

Usually it is difficult to distinguish between the two forms of learning.

Is Albert a case of classical or operant conditioning?

Is learned helplessness in dogs classical or operant conditioning?

There are two major differences between Pavlovian and operant conditioning:

1. In Pavlovian conditioning, the US is contingent on the CS, whereas in operant learning, a stimulus is contingent on behavior.

2. Pavlovian conditioning involves reflex behavior, such as the knee jerk reflex or salivation, whereas operant conditioning involves voluntary behavior, such as purchasing food.

In the Little Albert case, a Pavlovian explanation would be that the white rat became a CS for fear because it was paired with the loud bell to frighten Albert when he reached to touch the prior to being conditioned to fear the rat. An operant explanation is that Albert learned to fear the rat because he was punished for showing neutrality toward the rat.

Learned Helplessness – Overmeier and Seligman (1967) strapped a dog into a harness and presented a tone followed by a shock. The shock always followed the tone no matter what the dog did to try to escape the shock. Then the dog was placed in a different box that contained two chambers, one that elicited shock and one to where the dog could escape the shock. In spite of having an escape chamber, the dog did not try to escape the shock. It learned to be helpless.

A Pavlovian explanation would be that neither the tone nor the shock were contingent on the initial attempted escape behavior, leading to learned helplessness. An operant explanation is that everything that the dog did was punished, so the dog learned to do nothing.

Primary and Secondary Reinforcers

Primary Reinforcer – Innately reinforcing (no learning required)

Food, water, shelter, etc.

Secondary Reinforcer – Depends on being associated directly or indirectly with primary reinforcers (learning required, sometimes called conditioned reinforcers).

Praise, money, possessions, etc.

Advantages of Conditioned vs. Primary Reinforcers:

Conditional reinforcers have staying power.

Easier to reinforce behavior immediately with conditional reinforcers.

Less disruptive and are situation friendly (generalized reinforcer).

Secondary Reinforcer (Zimmerman, 1957) – A buzzer was sounded for two seconds before giving thirsty rats water. After repeated pairings, he placed a lever in the chamber. The rats soon learned that pressing the lever sounded the buzzer but did not dispense water. They continued to press the lever to sound the buzzer because the buzzer was a conditioned reinforcer. Money is a conditioned reinforcer because of its relationship with primary reinforcers.

Advantages of Conditional Reinforcers – Firstly, they have staying power and do not lose their reinforcement qualities as quickly as primary reinforcers. For example, there is only so much food you can give someone, but money is unlimited. Secondly, it is easier to reinforce behavior immediately using conditioned reinforcers because they are less disruptive. Eating and drinking take time, but praise is immediate. Finally, they can be used in a wide variety of situations. For example, food is only useful in one situation – when you are hungry, whereas money can be used in any situation and can be reserved for later use.

Generalized Reinforcer – A reinforcer that can be used in a wide variety of situations and is associated with many other reinforcers, such as money.

Disadvantages of Conditional Reinforcers – The biggest is that they occasionally must be paired with primary reinforcers to maintain their effectiveness, which is unlike primary reinforcers.

Shaping

What if the rat never pushes the lever inadvertently? What then?

Shaping involves reinforcing successive approximations of the final behavior (i.e. reinforce anything resembling the behavior at first, and then further reinforce closer approximations to the final behavior).

Laboratory Shaping

Shaping Undesirable Human Behavior

Shaping “In the Wild”

Laboratory Shaping – A rat dropped in the cage does not first realize that a lever in the cage will dispense all the food it wants as soon as it learns how to operate the lever. Once the rat is dropped in the cage and even begin the look in the direction of the lever, drop a food pellet. The rat will get the food pellet and then wander off. As soon as it looks again, drop another pellet. Eventually, the rat will move toward the lever and gets reinforced each time for every small movement. Soon, through successive approximations, the rat will press the lever itself to get the food.

Teaching the alphabet to a child is very similar. It eventually goes from being fragmented and incoherent in written from to smooth and legible through successive approximations.

Shaping Undesirable Human Behavior – Tantrums can be formed through shaping. Small amounts of whining results in bending to quiet down the child. When the same whining does not get the child what he/she wants the whining becomes more intense, which results in bending to more extreme whining. This can be shaped into full blown tantrums to achieve the desired goal.

Shaping “In the Wild” – To train their young to hunt in the wild, Otters will bring their kill to their young dead, then dying, then injured, and then eventually to the hunting ground.

Chaining

Shaping cannot account for all newly learned behavior.

Chaining – Learning a connected sequence of behavior, known as a chain. Chaining starts with task analysis.

Forward Chaining – Start by reinforcing the first link in the chain.

Backward Chaining – Start by reinforcing the last link in the chain.

It can be said that each link is a reinforcer of previous links, and primary reinforcement is received at the end.

Chaining – Skinner (1938) trained a rat named Plyny to pull a string that released a marble, then to pick the marble up and bring it to a tube, and then to drop the marble in the tube. Eventually, through chaining, the rat was able to complete the full task before receiving reinforcement.

Task Analysis – The process of breaking a chain down into its individual components. Once this is done, each component can be reinforced for being performed in the correct sequence. There are two forms of chaining: forward and backward chaining. Forward chaining involves starting at the first behavior in the sequence and linking it to the second until mastered, then adding a third until mastered, etc. Backward chaining involves starting with the last behavior and linking it to the second last, then to the third last, etc., mastering each addition to the chain before adding on.

Each link in a chain becomes a reinforcer.

REMEMBER, SHAPING IS USED IN CONCERT WITH CHAINING – SINGLE BEHAVIORS (SINGLE LINK IN CHAIN) MUST FIRST BE SHAPED, AND SHAPED BEHAVIORS ARE CHAINED TOGETEHER IN A SPECIFIC ORDER TO CREATE THE CHAIN INCLUSIVE.

Variables Affecting Reinforcement

Several variables affect operant learning:

Contingency

Contiguity

Reinforcer Characteristics

Task Characteristics

Deprivation Level

Contingency – Behavior is more likely to occur if it reliably leads to reinforcement (if lever pressing was independent of receiving food, the rat would stop pushing the lever).

Contiguity – The time between the behavior and the reinforcement, whereby the shorter the gap the faster the learning. For example, a 10 second delay led to the unsuccessful shaping of a bird to peck a lever (40 days), but a 1 second delay led to learning in 15-20 minutes. Immediate reinforcement of behavior is best. Why? Delay periods allow for other behavior to occur besides the target behavior, and if the inappropriate behavior is reinforced this will hinder learning.The effects of delay can be muted by signaling the delay period. In 1 experiment, rats were required to break a beam above their head, which led to food falling from the beam. Sometimes food was immediate and sometimes it was delayed. However, some rats received a signal (light) before the beginning of the delay period and some were not. The signal group did much better than the non-signal group; however, the immediate reinforcement group did the best. Some suggest that the marking hypothesis explains this finding, where the signal marks the behavior that occurred immediately before it was presented. However, a better explanation is that the light became a conditioned reinforcer.

Reinforcer Characteristics – The larger the reinforcer, the better the learning (however, the relationship is not linear – it is the law of diminishing returns). As well, there are individual differences with regard to preference for particular reinforcers.

Task Characteristics –Certain behaviors are easier to reinforce than others. Behaviors that involve internal smooth muscles and glands are difficult to reinforce, but behaviors based on skeletal muscle movements associated with voluntary movement are easier.

Deprivation Level – The effectiveness of primary reinforcers depends on whether the animal is deprived of such reinforcers. Animals who are full do not respond well to food as a reinforcer during learning. This is not the case for secondary reinforcers, such as money (you can be rich and still want money)

Other Variables – There are other variables that are important, such as prior experience with learning (tends to be a major difference between fast and slow learners), and the role of competing contingencies (when a behavior leads to opposite consequences depending on the situation.

Extinction of Reinforced Behavior

In operant learning, extinction refers to withholding consequences of behavior.

The overall effect is reduction of behavior; however, the behavior increases at first, which is known as extinction burst.

Another effect of operant extinction is behavior variability.

Extinction increases the frequency of emotional behavior, particularly aggression.

Spontaneous recovery will occur.

The possibility of resurgence of behavior during the extinction process of another unrelated behavior.

Behavior usually is acquired rapidly, yet it is extinguished slowly.

Can a reinforced behavior completely be extinguished?

Extinction burst is why tantrums get worse, but if you did not give in to these outbursts, the tantrums would cease.

Behavior variability occurs when the organism tries a new behavior (often a similar behavior) when a previously reinforced behavior is placed on extinction (tantrums may lead to “bargaining”). This can be used during shaping to get the animal to engage in a better approximation of the desired behavior (should be used delicately).

Extinction increases the frequency of emotional behavior, particularly aggression. Rats will bite the lever and even will bite another “innocent” rat. It also occurs in humans, such as wailing on a vending machine because it won’t release food, a tantruming child may throw things or break toys during extinction, etc.

Spontaneous recovery will occur, and the longer between the first and the second extinction session, the more pronounced spontaneous recovery will be. Gambling is a good example where the question is “why do gamblers go back to a machine that they know is not paying?” Why do you attempt to get food from a malfunctioning vending machine more than once?

The possibility of resurgence of behavior during the extinction process of another unrelated behavior, whereby the behavior that occurs usually is a previously reinforced behavior. For example, a child has been reinforced in the past for bargaining behavior in a store, which was extinguished. Now, the parent is in the process of extinguishing a new undesirable behavior, such as tantrums, but the bargaining behavior reappears. This may explain the concept regression, where a person has a tendency to return back to an infantile state/behavior unconsciously. If asking your wife to do something in a nice tone, which usually works, fails to lead to reinforcement, the man may regress back to a behavior that use to work on his mom, such as a tantrum.

How fast a behavior is extinguished depends on the number of times it was reinforced, the effort that the behavior requires, and the size/valence of the reinforcer of the behavior. However, learning is fast and extinction is slow.

Even when behavior has been extinguished it tends to occur more frequently than it did before it was learned. As well, a previously reinforced response that was extinguished is easier to learn the second time around.

Hull’s Drive Reduction Theory

Clark Hull believed that animals behave because of drives, which are motivational states.

All behavior is driven toward reinforcement. The drive reduces when reinforcement is received.

Does this apply to primary reinforcers only?

There are many reinforcers that do not reduce drives and are not related to primary reinforcers.

Next, we will discuss theories of reinforcement.

Hull believed that we behave because of drives (motivational state). Reinforcers reduce drives temporarily. The theory works well with primary reinforcers, but cannot explain many secondary reinforcers, or reinforcers that do not directly reduce physiological needs (praise, money, feedback). However, hull argued that secondary reinforcers derived their power from their association to primary reinforcers.

Many reinforcers do not clearly seem to be related to primary reinforcers. For example, babies will increase pacifier sucking behavior if it is contingent on improving the image of patterns on a screen. Why? Is this reducing a need? Male rats are reinforced by copulation even when they are interrupted before ejaculation. Is this reducing a need? Why do we prefer sweet water over regular water when the sweet water gives no additional nutritional advantages (actually is more harmful)? This is why Hull’s theory is poor.

Relative Value Theory

David Premack suggested that reinforcers are not just stimuli. They also are behaviors.

Certain behaviors are more probable to occur, making different behaviors have different values in a given situation.

The relative value of each behavior determines the reinforcing properties of the behavior.

The Premack Principle – strong behaviors strengthen weak behaviors.

The Premack principle is an empirically based theory as compared to Hull’s, but it is not without problems.

When a rat is given food for lever pressing is the food itself rewarding, or the act of eating (behavior). However, in certain situations, some behaviors are more likely to occur than others. A rat is more likely to eat than to press a lever. All behaviors have different values with regard to their reinforcing properties, and behaviors with higher values are more likely to be performed.

The Premack Principle – strong behaviors (high probability behaviors, such as eating) strengthen weak behaviors (lever pressing). How do we determine the values? Give the animal several options and the behavior that the animal engages in the most has the highest reinforcing value. For example, deprive a rat of water and make the administration of water contingent on running. The rat will quickly learn to run in a wheel. This suggests that running can be a reinforcer for water (have to drink water in order to run). This was successfully demonstrated.

The major problem with theory is explaining why many secondary reinforcers are reinforcing (such as saying “yes”).The biggest problem is that low probability behavior can be used to reinforce high probability behavior, which led to the next theory.

Response Deprivation Theory

Timberlake and Allison (1974) suggested that a behavior becomes reinforcing when the organism is prevented from engaging in it at a normal frequency.

All behaviors have a baseline of normal occurrence. When this baseline is decreased, the animal will engage in behaviors that re-establish baseline.

Any behavior that re-establishes baseline is considered reinforcing.

How is this different from Premack’s theory?

Secondary reinforcers and response deprivation theory.

Behaviors are reinforcing when an organism is prevented from doing them at baseline. Therefore, the relative value of a behavior is not important; what is important is engaging in the behavior at normal levels. For example, if you watch TV for 2 hours a day on average, and then TV watching is restricted to less time, you will be motivated to engage in behaviors that lead to the increase of TV watching to baseline levels. This can also apply to seemingly non-reinforcing behavior, such as eating peas.

The theory has a problem explaining praise reinforcers, such as yes, good, etc.

Two-Process Theory

Negative reinforcement also is referred to as escape-avoidance learning. How can we account for avoidance of aversives?

Two-process theory suggests that both Pavlovian and Operant learning are involved in avoidance.

According to the theory, there is no such thing as avoidance.

If the CS loses aversiveness, then avoidance should cease, but it does not, which is problematic.

The Sidman avoidance procedure is a big problem for two-process theory.

Anger’s (1963) proposal of time as a signal.

Negative reinforcement also is referred to as escape-avoidance learning. Place a dog in a shuttle box (there are 2 compartments with a door between them). Present a light and give the dog a shock 10 seconds later. The dog will always run to the other side to escape the shock. But eventually, the dog will learn to go to the other side without receiving shock, which is known as avoidance. This is an example of negative reinforcement, but why is the non-occurrence of something reinforcing? Explaining avoidance is a major problem, and 2 theories were create to account for avoidance: 1 and 2 process theories.

Two process theory – Avoidance involves Pavlovian and Operant conditioning. Regarding the dog example above, the escape portion is operant conditioning (negative reinforcement). However, the avoidance portion is classical conditioning, whereby turning off of the light becomes a CS for fear and the avoidance behavior is reinforced by escape from the shock compartment. Therefore, there is no such thing as avoidance; there only is escape (escape from shock and escape from the fear of shock).

Miller (1948) trained rats to escape shock where they were required to move from a white chamber to a black chamber of a shuttle box. Then he put them into the white compartment without giving a shock, but they still escaped. He then closed the door between compartments and made escape contingent on turning a wheel to open the door, which they learned even when shock was not presented. Then, he made escape contingent on pushing a lever, which they quickly learned even when shock did not occur. Apparently, the white compartment became a CS for fear leading to escape behavior, which must have been reinforcing.

Kamin (1963) trained rats to press a lever for food and then to escape shock via a tone signal in a shuttle box. Then returned the rat to the lever pressing, but sounded the shock tone to measure fear based on conditioned suppression. The greater number of avoidance trials experienced in the past, the less conditioned suppression or fear. If the CS becomes less fearful/aversive as avoidance training continues, what reinforces the avoidance behavior? Sidman (1966) avoidance procedure where a shock is given to a rat regularly where lever pressing delays the shock for 15 seconds. By pressing the lever regularly, the rat can avoid shock. Remember, nothing is signaling shock. If nothing is signaling shock (CS), what is the animal escaping? Both of these are detrimental to two-process theory. Anger (1963) believed that time was the signal. Even when time is factored out where it cannot be a signal, animals will learn to avoid aversives. All these problems led to the one-process theory.

One-Process theory

Escape and avoidance are reinforced by reducing aversives. which is operant learning only.

Herrnstein and Hineline (1966) support the one-process theory.

Avoidance can be extinguished by preventing the behavior and the consequence, which supports the theory.

Avoidance is only explained by operant conditioning. Herrstein and Hineline (1966) rats would receive shock every 7 seconds without pressing a lever and every 20 seconds if the lever was pressed repeatedly (these are averages, there is no guarantee that a lever press may lead to immediate shock). Animals pressed the lever to avoid shock. This proves the 1 process theory and proves that time does not become a CS for avoidance behavior.

If you force/constrain a dog to stay in a place it was shocked without the negative consequences ensuing, the dog will not try to escape on subsequent trials without constraints.