Showing posts with label no reward markers. Show all posts
Showing posts with label no reward markers. Show all posts

Monday, May 13, 2013

Shedd Animal Training Seminar: Aversives, Punishment, and You!

I don’t know why, but I always enjoy discussions on punishment. In ways, it feels like a “forbidden fruit.” I very rarely use punishment with my dog or my clients’ dogs, and if you try to discuss it- even theoretically- online, it can cause a lot of controversy. So my opportunities to talk about it are rare.

During the Shedd seminar, Ken talked about the advanced concepts of punishment, negative reinforcement, and aversive stimuli. These are three distinctly different concepts that are often confused, misused, and misunderstood. Still, the definitions are quite simple, and if you plan to use any of these techniques, you really do need to understand them.

An aversive stimulus is something that the animal wants to avoid. There is no definitive list of what makes something aversive; each animal will have different feelings about this. For example, some dogs hate being squirted in the face with water, but Maisy thinks it’s AWESOME.

A reinforcer is anything that increases the behavior it follows. Positive means something was added to make that behavior increase, while negative means something was removed. A negative reinforcer happens when something is removed, and as a result, a behavior increases in the future. This can happen for two reasons. First, the behavior may increase due to avoidance; an aversive isn’t actually applied, it’s simply threatened. The animal acts in order to prevent it from happening. Or, the behavior may be the result of escape. This happens when the aversive is actually applied and the removed with the desired behavior occurs. But either way, negative reinforcement is at play. It’s important to note that negative reinforcement can work and be both humane and effective if it’s done correctly.

A punisher is something that decreases the behavior it follows. This, too, can come in the positive or the negative variety. One way punishment can be used humanely is through deprivation; a reinforcer is withheld (negative) so that the animal will not perform the incorrect behavior again (punishment). Ken pointed out that this is why it’s so important to have multiple reinforcers available because this allows you to withhold certain reinforcers without depriving the animal of his full diet.

With that said, you really do need to know your audience when you use these terms. A trainer will punish a behavior; she wants a particular action to stop. But the public tends to punish the animal. That is, the punishment happens well after the fact, such as grounding a child for a bad report card or putting someone in jail for a crime they committed. In both cases, the actual behavior is so far removed from the consequence that it’s probably not being affected much.

So, while Ken does use punishment, he does not use it as the public understands it.

Ken talked about the use of conditioned punishers, as well. These are things that become aversive by association. Just as a clicker is a conditioned reinforcer because it predicts good things, there are also things that will predict bad things.

A delta signal, which is a warning to the animal that an aversive is about to be applied, can sometimes be used as a last chance to get things right. “Stop doing that or else,” it tells the animal. Your mom using your full name can be a delta signal; it tells you that you need to stop pulling your sister’s hair or face her wrath. The problem with deltas is that it can be very easy for the emotional trainer to escalate the use of punishment.

Ken also told us that a no reward marker acts as a punisher. This is the opposite of a bridge; it marks the moment when a behavior is wrong so the animal won’t do it again. These are typically quite mild, but can still cause frustration in the animal. So, while a skilled trainer can use no reward markers effectively and humanely, Ken thinks the potential for misuse is high.

I think my favorite part of this section was Ken’s discussion on how trainers use punishment versus how the public does. I appreciated the focus on behavior, not whether the animal is being “good” or “bad,” “cooperative” or “stubborn” (a word that always makes me crazy).

But what do you think? Anything intriguing here?

Thursday, December 2, 2010

Why I Love My Clicker

Crappy picture is crappy. But I didn't want to spend more
than five minutes shaping this behavior, so this is as good as it gets.


I’ve been thinking a lot about feedback this week. I’ve been thinking about Ian Dunbar’s approach to giving feedback. I’ve been thinking about the way I give feedback, and especially my own shortcomings in my rate of reinforcement. I’ve been thinking about how I react to criticism, and how my dog reacts to being told she’s wrong, too. And after thinking about it, I’ve come to the conclusion that I love my clicker.

Well, I already knew that. Most positive-reinforcement trainers use one. Even trainers who use collar corrections add in clicker work from time to time. It’s a great tool, which is why I was surprised that Ian seemed kind of anti-clicker during the seminar. It seemed to be about more than just the fact that he’s a pet dog person; his criticism of the clicker wasn’t about new students having difficulty with timing or misunderstanding the concept. Instead, he seemed to take issue with the feedback the clicker itself gives: impersonal, sterile, and devoid of emotion and instruction. But in a lot of ways, that’s exactly why I love it!

Is that weird? Maybe. After all, I do understand Ian’s point. All the clicker can do is say yes, you did that correctly. That’s it. It has one setting, one level, one message: yes. It can’t judge quality. It can’t say hey, that was even better than last time. Which is why I supplement the clicker with my voice sometimes, like I did in the chicken video. And if I’m going to do that, why should I bother using the clicker at all?

Because it’s different. Here’s the deal: our dogs hear our voices a lot. I talk to my husband, I talk on the phone, I talk to the cats, heck, I talk to myself. There are even toys that let us talk to our dogs when we’re gone! Of all the times that Maisy hears my voice, how often is it directed towards her? And even then, what percentage of that is meaningful communication versus me just chattering away at her because I love her? The vast majority of the time, my voice is simply background noise with little relevance to her life.

The clicker, on the other hand, is distinct. It’s easy to pick out of the sounds of day to day life because there isn’t anything else in the environment like it. More than that, though, it’s reliable and predictable. That sound always means something good is coming. It may only deliver one message, but that message is unambiguous, easy to understand, and worth paying attention to.

As a result, the clicker is able to get through to my dog much easier than my voice can. When Maisy is distracted or excited, she tends to tune me out. I can almost hear her sometimes: yeah, yeah, I know you’re talking, mom, but right now I’m concentrating so much on those chickens that I can’t be bothered to figure out if your words are for me or not. The beauty of the clicker is that she doesn’t need to think about it at all. She just knows that it’s for her.

Granted, this response happens because the clicker is a conditioned stimulus, not because it’s magic. Any sound can be conditioned the same way, including our voices. In fact, most clicker trainers have a verbal marker that they use, too. However, Lindsay Wood’s thesis found that it takes longer to train a behavior with a verbal marker than with a clicker.

In short, I’ve found that the clicker just relays the message better than my voice does. Of course, that’s still just one message. No matter how good it is at it, it can still only say one thing. I know that Ian really likes to say both yes and no, but honestly? I don’t. It’s not that I necessarily object to saying no- I understand that you need to inhibit behavior sometimes- but I prefer to focus on what’s going well.

I’m a sensitive person. As much as I learn from criticism, as much as I need it and want it, my ego is fragile. I do much better when I’m given positive feedback most of the time. The trainer that Maisy and I work with now is excellent at this, and when I fail, the positive feedback she’s given me previously is able to offset the current negative feedback. It helps me from taking it too personally, and the end result is that I feel generally confident in my abilities and I’m more willing to take risks, even if they might end in an error.

Maisy’s very sensitive, too, so I think she might feel the same way about negative feedback. But even if she doesn’t, and even if my next dog isn’t as soft as her, I know that I’m just not good at giving negative feedback. When I experimented with no reward markers earlier this year, I learned that when I have permission to say no, I quit saying yes. I become frustrated with my dog and with myself. Obviously, this does not help our training.

True, the clicker is just a tool, and like any tool, it has both good points and bad points. But for me, the clicker forces me to focus on the positive. It keeps me on track. It makes me look for success instead of dwell on failure. And most importantly, it builds up my confidence, Maisy’s confidence, and our confidence in one another.

And that’s why I love my clicker.

Sunday, November 28, 2010

Ian Dunbar Seminar: Providing Feedback to Our Dogs

One of Ian’s biggest criticisms of dog training today has to do with how we are providing feedback to our dogs. He believes that the vast majority of training is done with what he calls “Non-instructive Quantum Feedback,” while dogs would be much better off if they received “Instructive Analog Feedback.” It’s an interesting distinction, and today I’d like to spend some time explaining the difference, and why Ian’s so adamant that we switch over to the latter.

First, let’s break down what he means by these terms, starting with “Non-instructive Quantum Feedback.” “Non-instructive” means that the training simply provides consequences, either desirable (click and treat) or undesirable (collar correction), but does not explain why the action was correct or incorrect, nor to does it explain what the dog ought to do in the future. “Quantum” means that that the feedback can be counted in measured in some way. This makes it a simple response, and devoid of emotion.

“Instructive Analog Feedback,” on the other hand, is basically the opposite. “Instructive” means that we tell the dog what we want, either prior to the behavior (by using a lure), or after the dog does it wrong (by explaining what we wanted instead). “Analog” means that the feedback expresses value, which is difficult to measure. Ian does this by using his voice.

You will notice that this does not reference the four quadrants in any way. As I’ve previously written, Ian finds most of learning theory to be unnecessary to dog training, and he says the quadrants fall into that category. Instead of looking at feedback in one of those four ways, he sees it as binary: things either get better, or they get worse.

Instead of worrying about terminology (which is pointless anyway, because your intention and the dog’s perception may not match up), Ian says there are three types of feedback that people use. We reward the behavior somehow, we punish the behavior somehow, or we do nothing. Similarly, there are three ways of combining this feedback.

First, there’s the “old way,” of punishing all incorrect behaviors, but doing nothing when the dog gets it right. In this example, the dog often has no clue why he’s getting punished. It also relies on the use of painful consequences, such as a collar correction, which Ian says (and I agree) are unnecessary for dog training.

Then, there’s the “new way,” of rewarding the desired responses and ignoring the rest. Shaping falls into this category, and Ian isn’t a fan of it. He believes that dogs find shaping frustrating due to the lack of feedback when the trainer is silent. While I certainly believe that it would be frustrating to go “several minutes” without a click, if you find this happening during your shaping sessions, you’re doing it wrong. Shaping should split the criteria up into tiny increments; I once heard Kathy Sdao say that, during a shaping session, you should be clicking roughly every three to five seconds.

Ian also said that he doesn’t like shaping because if you click the wrong thing once, the dog will persist with that action, and you’ll be stuck there for long periods of time. Certainly you get what you click, but a single mis-click is pretty easy to overcome, especially if your criteria is split out well and your rate of reinforcement is high.

Finally, there’s the Ian’s preferred method: using both rewards and punishment. He believes using both helps the dog figure out the task faster. Keep in mind that he believes it’s possible to punish the dog without pain, so he is not combining collar corrections with treats. Instead, he’s using feedback that is instructive and analog, and he uses his voice to accomplish this.

Basically, he uses his language and his emotions to provide feedback to the dog. Not only can he tell the dog if he’s right or wrong, but also how well he did. This allows Ian to let the dog know whether the behavior was average or if it was truly exceptional. It also allows us to inform the dog how serious his misbehavior was, ranging from, “that wasn’t quite it” to “holy crap, that was dangerous!” Ian says this allows the dog to receive far more information about his actions than a simple click and treat, or a buzz and shock from a collar.

Personally, when I’m teaching Maisy new tasks, I use primarily shaping. I do not tell her when she’s getting it “wrong”- but then, I don’t think a dog can be wrong when shaping anyway, since the whole point is that the dog is supposed to offer behaviors and you choose the ones you want to work with. I have tried using no reward markers- when you tell the dog they’re on the wrong track- but I found that they cause Maisy to give up. However, I do add verbal feedback when Maisy does something amazing. Click- jackpot- and lots of praise. Although Ian seems to believe clicker trainers don’t do this- and maybe purists don’t, I guess I don't know for sure- I don’t see why I can’t use my voice in conjunction with clicker training.

As for what I do when Maisy performs a known behavior incorrectly, well, that’s a discussion for another day. Ian talked a lot about how he approaches this: he uses what he formerly called an “instructive reprimand.” However, once he found out that people interpreted “reprimand” to mean something harsh, like yelling, he renamed the method “repetitive reinstruction as negative reinforcement.” This is a very interesting method, and it deserves its own post. I’ll do that very soon.

In the meantime, I’d love to hear about how you give your dog feedback, especially when teaching new behaviors. I know there are many ways to train, and I can’t wait to hear how different people approach this!

Monday, April 26, 2010

Suzanne Clothier Seminar: Wrap Up

Wow, who knew that a two day seminar could inspire over a month’s worth of blogging? I want to thank everyone who commented on these blog posts. The discussions we’ve had over the past month have really helped me think about the things Suzanne said in a far more sophisticated manner than I could have alone.

I just want to touch on a few of the highlights from those conversations today. If you haven’t, I encourage you to go back and read the comment threads. There are a lot of smart, dedicated, and talented people in there sharing differing perspectives. Although we dog trainers will probably never agree 100%, it’s nice to consider other ideas, either to refine our own thoughts, or to strengthen our positions.

For me, I have found that the seminar and resulting discussions have strengthened my commitment to positive training, even if the term is a bit of a misnomer. As everyone noted, it is impossible to use solely positive reinforcement. I strive to teach enough foundation skills that I rarely need to stray from that principle of operant conditioning.

Still, there are times when a consequence is needed for a less-than-desirable behavior. The challenge is to find such a consequence which is neither physically painful nor which causes excessive emotional stress. Of course, there is the challenge of defining how much stress is too much, but I’m afraid I have yet to figure that one out. So far, it seems to be a matter of knowing the dog well enough to be able to stop while we’re ahead, but that’s a rather ambiguous answer, and one which is undoubtedly frustrating for the less experienced trainers out there.

I have decided that for my dog, the best way to deal with unwanted behaviors is to use Premack’s Principle. I’ll admit, while I understand the principle intellectually, I don’t quite get why it works so well. At any rate, I’ve had some amazing results with Premack, and so have others.

When it comes to our “silly tricks”- my name for competition behaviors which really only matter because I have a goofy hobby, and not because they’re vital life skills- the consequences for an incorrect response is generally removing the reinforcement or doing a time out. Time outs work well; Maisy loves to train, and she loves to spend time with me. Removing my attention for a short period of time is a far more effective punisher than withholding a food treat.

I do occasionally use mild verbal corrections, but Maisy is so sensitive that I have to be careful with using these. I try to avoid them, as well as no reward markers because they tend to frustrate both of us. Similarly, I use some pressure/release techniques with her, such as body blocks or light physical pressure, but I have to be careful with these, too- she’s incredibly sensitive to physical touch. In fact, I once tried using a body wrap on her, which are widely promoted for reducing anxiety. It did not go over well, and in fact, actually caused more anxiety. Although both verbal corrections and physical prompts can be useful tools which fall on the more positive end of the “consequence spectrum,” they are things which I must use sparingly with my dog.

Which brings me to my favorite part of the seminar: Suzanne’s repeated insistence that we view all dogs as individuals. I love that she says training is humane only when we check in with the dog regularly in order to get his perspective. Can you do this? Is this okay with you? How can I help you? These questions focus on building up the relationship in the name of training, and I’ve always said that training is only about the relationship between me and my dog anyway.

Finally, I think the biggest benefit I got from the seminar was learning to give Maisy the information she needs to be successful. Her statement that dogs look to their people for clues on how they should react really encouraged me to look at what role I play in Maisy’s reactivity. I’ve always known that Maisy is sensitive to my moods and reactions, but being forced to confront that reality at the seminar has really improved my awareness of my own body. Making a conscious effort to remain calm and confident in the face of triggers has gone a long way in soothing Maisy’s fears.

All in all, that weekend was one well spent. I really think I grew a lot as a trainer as a result of the things Suzanne said, and again, I really appreciate each and every one of your comments.

Wednesday, March 10, 2010

Asking the Wrong Questions

I’ve spent a lot of time over the last week thinking about no reward markers and keep going signals. In fact, I’ve thought so much about it, that I just had to post about it again.

In the comments to my last post on this topic, the point came up that there is a huge difference between shaping and competition. This is absolutely true. Shaping is about teaching a new skill, while competition is about testing a skill which is, presumably, under stimulus control. This means that you can’t really compare how a dog interprets silence from the learning stage to the performing stage; they’re two completely different contexts.

More than that, though, I realized I was also asking the wrong question entirely. When it comes to shaping, the question should not be How does my dog interpret silence? Instead, the question should be Why is there silence at all?

Think about that for a moment.

Now think about your last shaping session with your dog. How much silence was there? And why was there that much silence? For my last session, there was about thirty seconds of silence. Why was there that much silence? Well, because Maisy didn’t meet my criteria, of course.

But is that true?

My job as a clicker trainer is two-fold: split the task down into many small steps, and give a high rate of reinforcement when my criteria is met. These two things are interrelated. If a task is properly broken down into small, achievable steps, your rate of reinforcement will naturally be quite high. Likewise, the inverse is true: if you lump the steps together by setting the criteria too high, it will take your dog longer to figure it out, and thus your rate of reinforcement will be lower.

So why was there that much silence? Because I failed to do my job as a trainer. I lumped when I should have split.

Clicker training is difficult to master. To be a truly efficient trainer, you need to not only be able to split the task up into small steps, but you also need to be able to analyze your dog’s response, assess whether that means your criteria is too high, too low, or just right, and then adjust that criteria… and you need to be able to do all of that in a matter of seconds!

Thankfully, clicker training is also easy to learn. Even if you never move beyond the basic "click the behavior you like and give your dog a treat" stage, your dog will still learn. That's what I love about clicker training: regardless of your skill level, it has something to offer to everyone.

Thursday, March 4, 2010

Clicker Theory: No Reward Markers and Keep Going Signals

I apologize in advance to any readers who are not familiar with clicker training, or who are just beginning to learn about the learning theory behind it, as today’s post concerns more sophisticated clicker concepts.

About a week ago, someone on a mailing list I belong to posed a very interesting question: If the click means “yes,” then what does no click mean?

The poster, a teacher, mentioned that when she has her students play the“clicker game,” in class, they initially learn faster if they receive feedback for both yes, that’s what I want you to do, and no, you’re going in the wrong direction. In other words, a no reward marker. She went on to say that once her students understood the game, they learned that the absence of a click or verbal marker was basically the same thing as being told no. Once they figured that out, they could figure out the task just as quickly with only the positive marker.

She wondered: do our dogs understand the absence of a click the same way? Do they interpret silence as “no”? If so, why do they keep working in trial settings, where they receive neither clicks nor encouraging verbal feedback? Wouldn’t the silence inherit in a trial tell our dogs that they are doing it wrong? If so, this would have dire consequences on our performances.

The general response was that silence should not- cannot- imply that the dog made an error. Instead, we must teach our dogs that silence is a keep going signal- that they are on the right track, and that if they keep up with what they are doing, they will earn reinforcement. That is the only way that our trial performances will hold up.

So, if silence means “keep going,” then how do we tell our dogs they’re going off track? As Clicker Trainers, we don’t use corrections (defined here as anything that causes the dog pain or stress) to tell the dog they’re wrong. The logical response would be the use of a no reward marker- an emotionally neutral way of saying no, try something different… but some people on the list argued that this would actually slow learning down.

I disagreed. I shared with the group that when Maisy begins to get off track during a shaping session, I tell her “Nope! Try Again!” in a cheery voice. I wrote that I felt my dog learns faster this way, but that even if she doesn’t, it helps me feel better to be giving the feedback.

Still, in light of the conversation, I decided that I would test my theory, so I sat down with Maisy to work on a shaping project. First, I just worked with her like normal, not really thinking about what I say or when I say it. Although I did say “Nope! Try Again!” perhaps three or four times in the course of five minutes, I found that I said it more as conversation and less as information. Interestingly, I discovered that I was saying it at times when we were in the midst of a long period of silence. That “nope!” served to fill the silence until she finally got the click for doing what I wanted.

Next, I worked with her, but remained silent. I didn’t speak; I simply clicked or didn’t. Maisy continued on, doing well until we hit one of those long periods of silence. She kept trying things, but after about thirty seconds of neither a click nor a “nope!”, she laid down and looked at me as if she wasn’t sure what she was supposed to do.

Finally, I tried using the no reward marker more regularly. We continued shaping, but I tried to think in terms of right and wrong. I clicked when she got it right, and said “Nope!” when she got it wrong. This led to rapid-fire clicks and “nopes,” and after she got three “nopes” in the space of about ten seconds, Maisy again laid down with her chin on the floor. This time, though, I had to encourage her quite a bit to re-engage with the shaping game. But when she again got several more “nopes,” she laid down and refused to play any more.

I began to feel frustrated; this is not how it’s supposed to work! She’s supposed to want to play! My frustration came out in my voice, and I began to tell her to get up with an edgy tone. When she didn’t, my feelings of frustration gave way to anger. Since I didn’t want to take that out on her, I ended the session to evaluate what had just happened.

The first thing that I decided was that I was wrong: Maisy does not learn faster with a no reward marker. In fact, she gave up so quickly, and was so difficult to persuade to re-engage with the task, that I believe she found it punishing. True, she also gave up when the silence went on too long in the second scenario, but she worked approximately three times longer, and was much more willing to re-engage when I asked. As a result, I think she found the lack of any feedback confusing, but not aversive.

Still, I concluded that the complete lack of any kind of feedback was also not the best way to help Maisy learn. Instead, her learning is most efficient when she gets lots of reinforcement over a short period of time. This means my job is to break the shaping task at hand down into as many pieces as possible so it is easier for her to progress through each step of the task. However, sometimes it is difficult to figure out how to break a task down any further. As a result, if I cannot figure out how to make the task easier, and if it’s been fifteen to twenty seconds without a click, I need to give Maisy a “gimme” click- reverting to the previous level of criteria for a few moments before trying the higher criteria again.

I also suspect that my initial use of “nope!” wasn’t actually serving as a no reward marker. Given the way Maisy responded, I think it actually served the purpose of a keep going signal for her. This means that for tasks that haven’t had a sufficient amount of duration built in yet, she depends on verbal encouragement to know that she’s doing what I want. (Interestingly, though completely off topic, I haven’t been very good at building duration past 30 seconds or so, which was Maisy’s threshold for silence during these tests. It makes me wonder if my inability to build more duration in her behaviors is due to her threshold, or if she’s developed that threshold because I have neglected to put in the work necessary to build more duration. On second though, I’m pretty sure I know the answer to that.)

Finally, and perhaps most importantly, I learned that I don’t like it when I have to tell Maisy she’s wrong. I became frustrated and then angry as she continued to fail, even though that “failure” was behaviorally no different than when we did silence only, or when I used the keep going signal. Maisy was going about the shaping task in the exact same way in each scenario. She wasn’t any more wrong when I told her she was than when I didn’t. In other words: focusing on the wrong behavior rather than the right one changed the way I viewed and felt about the training session, and it took all of the fun and joy out of playing the shaping game with Maisy.

In the end, doing this experiment not only taught me that my initial supposition was wrong, but it also reaffirmed my commitment to positive training. Focusing on what I want her to do helps Maisy learn faster, but it also makes us both feel better.