Showing posts with label reinforcement schedules. Show all posts
Showing posts with label reinforcement schedules. Show all posts

Friday, August 9, 2013

CPDT Study Session #2: Schedules of Reinforcement

I’m in the middle of the two week period of time I’ve set aside to study for the learning theory part of the exam. I actually haven’t done much reading yet, both because I’ve been busy and because I’m pretty confident about my knowledge in this section. One thing I did want to firm up was my understanding of basic schedules of reinforcement. These schedules specify the timing and frequency of reinforcement, and each type can be useful in the right situation.

Continuous Reinforcement Schedules (CRF)

Most training starts here, with the continuous rate of reinforcement. This means that every time the dog does the behavior, he gets reinforced. It works best during the teaching phase, and it helps establish a strong contingency between the behavior and the reinforcer.

If you use a continuous reinforcement schedule, keep in mind that these behaviors are quite susceptible to “extinction,” which means that if you stop reinforcing the behavior, the dog is going to stop the behavior. Since it can be difficult to be sure that you reinforce every instance of a behavior, this schedule is a bit impractical. This is why most trainers switch to some kind of variable reinforcement schedule, but it is possible to use a continuous rate for the life of an animal (indeed, it’s what the Baileys- arguably some of the best animal trainers of the 20th century- used most of the time).

Partial (or Intermittent) Schedules (PRF)

There are several types of partial (sometimes called intermittent) reinforcement schedules. Although each type is distinct from the others, they do have several things in common. These are used when a continuous schedule is simply too cumbersome, for whatever reason. They are more resistant to extinction, and they typically feel more “natural” to people. You do need to be cautious that you don’t “thin” the schedule too quickly as this will cause “ratio strain” and degrade the quality of the behavior.

Fixed Ratio (FR)

A fixed ratio is when the reinforcer is given after a certain number of behaviors. The number after the abbreviation informs you how many behaviors need to be done before reinforcement is earned. For example, an FR5 means the dog must do five sits (or whatever) before receiving his treat.

Fixed ratios will produce high, steady rates of responding due to their systematic, consistent, and predictable nature. That said, fixed ratios also have a “post reinforcement pause” where the dog will briefly stop doing the behavior immediately after being reinforced. Their response time will increase as they approach the next opportunity for reinforcement. If your ratio is very high (such as an FR400), the post reinforcement pause will be longer.

Variable Ratio (VR)

In a variable ratio, the frequency of treats given is variable from trial to trial and should happen after an unpredictable number of times. It’s typically done around an average number of times. For example, a VR4 would mean that the treats are given approximately 1 out of 4 responses. During a series of behaviors, the treat may be given on the 2nd repetition, the 6th repetition, and then the 4th repetition.

Variable ratios yield high, steady rates of responding, and there is a much lower rate of post response pauses. This schedule is also more resistant to extinction and useful for fading out a fixed ratio schedule. That said, a truly variable ratio is difficult to achieve as we humans tend to be pattern dependent.

Fixed Interval (FI)

In a fixed interval, reinforcement is given after a certain period of time. An FI5 would indicate that reinforcement is given for the first correct behavior after 5 seconds (or minutes, depending) has passed since the last reinforcement.

Interval schedules (both fixed and variable) are great for teaching duration behaviors. A fixed interval is prone to extinction, though, and has a pronounced post reinforcement pause. In this case, the pause is “scallop-shaped;” the behavior levels off in the first bit of time, and then increases in frequency as the time for reinforcement comes due. This is similar to a student checking the clock more frequently when class is almost over.

Variable Interval (VI)

In this schedule, reinforcement is given on an average amount of time, which means the first correct behavior after an unpredictable amount of time has passed is reinforced. Like the variable ratio, a VI4 would mean that the reinforcement happens approximately every 4 seconds (minutes, etc.), but that the amount of time elapsed will change from trial to trial.

This schedule produces a slow, steady rate of responding, although you don’t tend to get a particularly high rate of behavior. It has good resistance to extinction, making it particularly good for fading out a fixed interval schedule. Like the variable ratio, it can be difficult to be truly unpredictable.

When I get around to it, I’ll post about the differential reinforcement schedules. There are quite a few of these, and they are arguably more interesting than these more basic schedules. But for now- what are you guys studying?

Monday, May 6, 2013

Shedd Animal Training Seminar: Advanced Concepts in Reinforcement

Okay, gang, I’m back with the Shedd Animal Training Seminar recaps. It’s been awhile, but thankfully I left off at a pretty good breaking point because we’ve come to the section on advanced concepts.

Ken defined advanced concepts as those that require experience in order to apply them. This is any training that ventures past the basics of “reward behaviors you want and ignore the ones you don’t.” You know you’re ready to start dabbling in some of these concepts when you understand training theory well enough to know when to ask for help (seriously. All good trainers get in over their heads sometimes) and you have some good mechanical skills (able to use a marker with good time, able to deliver reinforcers efficiently and effectively).

That said, just because YOU are ready to use an advanced concept does not mean that your animal (or your human client) is ready for the concept. So you also need to know when it’s appropriate to use one of these concepts, and when to stick with the basics.

A great example of this is the concept of defining criteria for a behavior. In the early stages, we think of behavior as a black-or-white kind of thing: either the behavior was 100% correct, or it was wrong. Except… there IS a gray area in training. This happens fairly often when a behavior is still in training, especially when you’re shaping a behavior with a series of approximations. Sometimes the animal gives you something you weren’t looking for or expecting, and you need to make a quick judgment call about whether or not to mark it.

With that out of the way, let’s talk a bit about when Ken considers reinforcement to be an advanced concept.

Being sprayed by a water bottle is a secondary reinforcer for this dolphin.

One situation in which using reinforcement requires an experienced trainer is when a secondary reinforcer is being used. Also called a conditioned reinforcer, this is something that the animal is taught to value. The most common example is a clicker or marker, but it’s anything that any animal will accept as a reinforcer. Secondary reinforcers can be indispensable when an animal is sick and is refusing to eat but you need to give them medications or reward them for a behavior.

Ken notes that your relationship to the animal is critical when you’re using a secondary reinforcer; while a kiss from your significant other may be welcomed, a kiss from your boss probably won’t be. For a more in depth discussion on secondary reinforcers, please see this post. http://reactivechampion.blogspot.com/2011/08/ken-ramirez-seminar-non-food.html

Another reinforcement technique that Ken considers to be an advanced concept is the use of variable reinforcement. Ken likes to look at reinforcement schedules simply. Instead of all the technical terms like CRF, FI, FR, VI, VR, etc., he tends to see them as either continuous and consistent or variable and intermittent. Of course, he readily agrees that understanding the technical terms can be helpful, but said that most of the time, it really isn’t necessary in most situations.

Variable reinforcement happens when an animal does not get a reinforcer for each and every behavior. It’s often used in training because it makes a behavior more resistant to extinction. This allows you to have the animal do a number of behaviors for only one reinforcer. However, it does need to be carefully introduced or it can lead to frustration in your animal.

Although there are many ways to introduce a variable schedule of reinforcement, Ken shared how the Shedd staff do it. First, every new trainer AND every new animal begins with a continuous, fixed schedule of reinforcement. They will provide a variety in the types of reinforcers, though. Then, they condition and establish secondary reinforcers (see the post linked above for more details on this). Next, start using your secondary reinforcers so that they are not always followed by a primary reinforcer. Finally, use other well-established behaviors as a reinforcer. This entire process generally takes four to six weeks with an experience trainer AND an experienced animal. With a naïve trainer and animal combo, it can take several years.

There is one more advanced concept in regards to reinforcement that Ken discussed: negative reinforcement. However, I decided it makes more sense to present it with the seminar summary on aversives and punishment. Keep an eye out for the next installment in the Shedd Animal Training Seminar series!

Tuesday, January 17, 2012

The Pleasure of Anticipation

Last spring, I wrote about how cues can be reinforcing for dogs. If the cue predicts a good outcome (a click and treat, for example), then the dog will find the cue exciting. More talented trainers than I have taken advantage of that by reinforcing a dog’s response with another cue.

Some readers met this with skepticism. Maybe my explanations made sense, maybe they didn’t, but let’s be honest: logic and anecdotes alone are not always convincing. That’s fine; I don’t expect everyone to agree with me, and in fact, I would find that rather boring. But when one of those skeptics found this hour-long lecture, she remembered my post and emailed me.

The lecture, given by neurobiologist Robert Sapolsky, explored what makes humans unique. His entire talk is fabulous, and I urge you to watch the entire thing. Personally, I really enjoyed his discussion of how language affects our perceptions of others because of the insular cortex, but what’s relevant today is what he shares about dopamine (starts about 30 minutes in).

Throw it... throooow iiiiiitttttttt.....
Dopamine is a neurotransmitter that helps control the brain’s pleasure and reward centers. For many years, it was believed that when someone (human or animal, it doesn’t matter- dopamine is present in all mammalian brains) received a reward, their brain would release dopamine. In turn, this would result in a pleasurable feeling.

However, when scientists actually studied what was going on, they found something very different. Sapolsky described an experiment in which chimps could receive a food reward if they press a lever when a light turns on. The dopamine levels in the chimps’ brains increased not when they completed the task, but rather when the light went on.

In other words, what the chimps found pleasurable was the opportunity to receive a reward, not the reward itself. After they pressed the level, their brain quit releasing dopamine, even before they received the reward. Anticipating the reward was better than the reward itself.

That light signified an opportunity to receive a reward; press the lever now, it said, and you will be reinforced. This is exactly what we do in dog training. I say “sit,” and if my dog does, she’ll get a treat. So the light was acting as a cue. The study Sapolsky cited says that it was the cue that made dopamine levels rise, which means that my dog will feel good when I say “sit,” not when I give her the treat. The cue is reinforcing.

I suspect that clickers work the same way, although Sapolsky didn’t address that directly. He did, however, say that dopamine is about the anticipation of the reward, not the reward itself. If cues can cause that anticipation, it seems that a sound could, too. Can a click cause dopamine levels to increase because the dog is now expecting to receive his reward? I don’t see why not.

What scientists found even more remarkable, however, was that when the food was given in response to the correct behavior only half the time, the chimps’ dopamine levels went through the roof. This wasn’t exactly surprising to me; dog trainers often talk about how a variable schedule of reinforcement creates stronger, more durable behaviors than when the dog gets a treat for every correct behavior. B.F. Skinner and his students proved that over and over again in the lab, although of course they couldn’t know that it was the result of dopamine. As Sapolsky put it, “maybe is addictive like nothing else.”

Finally, the scientists also found that if they blocked dopamine production in the chimps’ brains, when the light came on, the chimps didn’t care. Instead of eagerly pressing the lever, they sort of shrugged it off. The chimps knew they’d get a reward if they did, but they just didn’t seem to care. Could this be a possible explanation for why a dog doesn’t respond to a cue? Maybe. But I'd point out that there are many, many other reasons dogs don’t perform a behavior, and most of them are probably more logical. Still, it is fun to think about.

I found all of this really interesting. Not only did it support the concept of cues being reinforcing- something I find pretty fascinating in and of itself- but it also suggests that there is more at play in clicker training than just the food. In fact, it would seem that anticipation is what's truly powerful, an idea I find amusing since trainers often get upset when their dogs anticipate what's coming next.

To be fair, having the dog act before we ask them to can be a problem. Still, is that indicative of a corresponding spike in dopamine? And if so... how can we use this to our advantage? What can we do to harness our dog's natural brain chemistry to create a more favorable training outcome? I'll admit, I don't have an answer here, so I turn it over to you: have you ever used the power of anticipation to your advantage? And if so, how?

Tuesday, November 23, 2010

Ian Dunbar Seminar: Is Learning Theory Useful for Dog Trainers?

Ian Dunbar is not a heretic- he wants you to know that. Operant conditioning is real. It’s been proven. The experiments behind the theory were repeated thousands of times, and they were validated. It’s great science.

It’s just not useful for dog trainers. Or at least, not most of it. Ian believes that only about 10% of learning theory applies to us because learning theory was laboratory-generated. That is, the experiments were implemented, monitored and controlled by computers, carried out using Skinner boxes, and the animals used were “simple” and had “few interests.” In contrast, we are humans, and we train real dogs in the real world.

If we choose to use learning theory in training, then we must learn to train like a computer. That’s not necessarily bad- computers have some pretty good traits. They are tireless. They are completely consistent in both monitoring the trainee’s behavior and in providing feedback. This allows them to have very clear criteria. By contrast, we humans often have unrealistic or unclear criteria, and are inconsistent in our observations and feedback. So, why wouldn’t we want to train like a computer?

Well, computers as trainers have some drawbacks. They can not qualitatively assess an animal’s performance- they can’t see cute or flashy behaviors and train that into the final product. While they can give feedback, they cannot give instructive feedback. All a computer can do is say yes or no- provide a click and treat or a buzz and shock. They cannot explain why the animal was wrong or what he should do instead, and they cannot explain how important or urgent compliance is. Humans can.

Then there is the matter of reinforcement schedules. Ian identified seven reinforcement schedules: continuous, fixed interval, fixed ratio, variable interval, variable ratio, random, and differential. Ian explained that six of these seven schedules will maintain a behavior, but only one will improve behavior. The one that will improve behavior- differential reinforcement- is the one that computers cannot use. (Personally, I disagree. Computers may not be as good at it as we are, but there is no reason a computer couldn’t reward faster responses. In fact, they might be better at that than I am- I do not have a stopwatch in my head.)

Because we humans cannot be as consistent as computers, and because computers cannot provide instructive feedback the way we can, Ian sees no need for us to try to emulate computers. He finds this to be especially true because most dog owners don’t need nor want the precision that comes about from training like a computer. As a result, he really doesn’t have any use for the vast majority of learning theory.

So what does Ian like? Thorndike’s Law of Effect, which more or less says you should reward the good stuff and punish the bad stuff. Ian says this is simple, elegant, and pure. It doesn’t get into complicated and confusing types of rewards or punishment which cause endless arguments on the internet. Thorndike tells it like it is.

Again, Ian’s orientation as a trainer of pet dogs is obvious. The average dog owner doesn’t care about precision, and doesn’t have the consistency or patience needed to sort through the various quadrants and schedules, so I understand why Ian thinks we should avoid discussing learning theory with clients. We need to quit worrying about the science and terminology and just train. We should help them, not confuse them. It’s hard to argue with that.

Still, as a dog geek, I struggle with this idea. Personally, I enjoy understanding the science behind what I’m doing. Ian said it himself: learning theory is valid. I like thinking about what I’m going to do. I love planning my sessions. I also think it’s fun to take data and evaluate what I’ve done with the goal of doing better next time.

As a competitor, I want precision. I enjoy the challenge of being consistent enough to get amazing results. I strive to be as clear as possible in my criteria. In many ways, I do try to train like a computer, and I don’t think that limits me. I enjoy pairing clicks with not only treats but also heart-felt praise when my dog does something exceptional. I see no reason to have to choose between computer or human. That’s sort of the beauty of being human, after all: I can think outside of the box and combine the best of both approaches.

I know I’m not the normal dog owner. I spent Halloween weekend at Ian’s seminar, and I’m spending hours writing up my notes for this blog, after all. I would ask all of you if you’re normal dog owners, but I suspect I know the answer to that. You are reading this, after all.

Instead, go ahead and analyze what Ian said. Tell me how it makes sense, and then how it confuses you. Tell me how you use, or don’t use, learning theory with your dog. I know you’ll have lots to say!