Title: Part 4 Learning from Feedback with Tools/Code Authors/Date: Akanksha Bhardwaj, Azalia Mirhoseini — Stanford CS329A Video ID: Lxh9RF5S-K0 | URL: https://www.youtube.com/watch?v=Lxh9RF5S-K0 Playlist: https://www.youtube.com/playlist?list=PLangBM27OtEA ------------------------------------------------------------------------ 0:05 In the first lecture, that be covered 0:07 is that large language models are 0:09 really good at solving natural language processing problems, 0:12 and I think most of you are very familiar with using them 0:15 in chatbot scenarios. 0:17 But to make them useful in real-world tasks, 0:19 they do need to be able to interact 0:21 with real-world environments, tools, and code, 0:25 and learn from what they interact with. 0:28 So you already saw some examples of test-time compute 0:32 and robust verification that Azalea covered in the last two 0:35 lectures. 0:35 Now, we'll build a little bit more on top of that. 0:38 So today's lecture, we will focus on three different papers. 0:42 The first paper will start with the notion of tool calling, 0:46 and how can we get the models to go beyond just reasoning 0:50 to also take into account actions via tool calls? 0:53 And the paper that we'll cover here 0:55 is called ReAct, which actually was one of the first 0:58 that went around to combine reasoning and acting in language 1:01 models. 1:03 The second paper that we'll look into 1:05 is RLEF or grounding code LLMs in execution feedback. 1:09 What this will focus on is how do you actually 1:13 do code generation while using execution feedback from being 1:17 able to run the code or run the tests of the code. 1:20 And the third thing that we'll look into 1:22 is constitutional AI, where we use the model critiquing itself 1:27 as a feedback loop to improve the model. 1:29 So all three of those techniques are ways for the model 1:32 to improve itself. 1:34 But what differs in each of these techniques 1:36 is, where is the feedback coming from? 1:38 It's some form of interaction with either the environment 1:41 or some form of critique mechanism or grounding 1:44 using what we know should be the right answer. 1:48 So let's start with tool calling as the first abstraction. 1:56 So what ReAct really tried to focus on 2:00 was this notion of when humans want to go and do some action, 2:03 they think about it first. 2:05 For example, if you decide that you want to come take a class, 2:09 you think about how that might actually help you. 2:13 And what this roughly means is that, once you 2:17 have thought about it, then you will have a reason to act. 2:20 And then based on that action, you 2:22 might observe that you have a new observation 2:24 about the environment. 2:25 So once you come to the class, you 2:27 might decide you like the class or you don't like the class. 2:29 And based on that information, you come up with new reasoning. 2:33 Maybe you'll go do the homework, and then that 2:35 will make you decide something else. 2:36 So there is this loop between thinking and action 2:39 that helps us function as humans. 2:42 A similar loop can be very useful in terms of refining 2:46 what language models generate. 2:48 And typically, when you look at language models, 2:52 they have struggled to combine reasoning and acting 2:55 by themselves as isolated processes. 2:57 In today's models, the reasoning has become innate 3:00 because thinking has become part of the models. 3:02 But if you look at the history of the LLMs 3:04 in the last couple of years, this 3:06 was not something that came out of the box. 3:09 And one of the challenges of why it was hard for models 3:12 to combine reasoning and acting is because they typically 3:15 will hallucinate. 3:16 They will have no context of what is actually 3:18 happening in their environment. 3:20 So, for example, if you were to type 3:21 into a model like Gemini or ChatGPT 3:24 or Claude, what is the temperature today? 3:30 Where is something happening today? 3:31 It has no information about the current events, 3:34 unless it actually goes and does a search call 3:37 and then fetches that information. 3:39 So it has no notion of grounding in the real world. 3:42 And as a result of that, it cannot go and take actions, 3:46 and then actually combine that back into reasoning. 3:49 So ReAct was one of the first abstractions 3:51 that tried to combine it in a way that would be effective. 3:54 And what you see today's models do 3:57 is actually go beyond ReAct to do this in the chain of thought 4:01 itself. 4:02 So let's take a look at what we covered in lecture 1. 4:05 So if you look at chain of thought by itself, 4:08 chain of thought already gives the model 4:10 some sense of reasoning. 4:11 So if you gave it a math problem, 4:13 the model will show the steps, the intermediate steps, it 4:16 takes to get to the answer, but these steps may or may not 4:19 be grounded because they are based on models internal state. 4:22 They have no feedback from the external world. 4:25 And then there is a second class of LLMs 4:27 that were developed, for example, WebGPT, 4:29 which go and interact with, say, web browsers. 4:31 And what they're basically learning 4:33 is how to interact with web browsers, which 4:36 is a form of a tool. 4:37 So they don't necessarily have a reasoning process, 4:40 but they're basically learning how 4:41 to interact with web browsers. 4:43 So you basically can fine-tune on traces 4:46 that are interaction traces with web browsers, 4:49 and then what do humans prefer? 4:51 But as a result of those two different things, 4:54 one is focused more on reasoning and one 4:55 is focused more on actions. 4:58 How do we combine these two paradigms 5:00 where we can both reason and act in a useful way? 5:03 A very simple abstraction that ReAct came up with 5:06 was just using prompting. 5:08 You can get the model to generate verbal reasoning 5:10 traces. 5:11 So first, you say, OK, I want the model to go do this task 5:15 and ask the model to think, say, step by step about it, 5:18 and then it comes back with its reasoning trace. 5:20 You append that in the prompt, and then you 5:23 say, OK, what action should be taken based on that? 5:25 And then you go take an action. 5:27 This really helps the model improve performance 5:30 where grounding matters. 5:32 So HotpotQA, for example, is a question-answering task. 5:34 Then FEVER is a fact-checking task. 5:36 So these require real-world knowledge. 5:38 They require knowledge about what is happening in the world. 5:41 And then WebShop, for example, is a task 5:44 where you're supposed to ask the model to go 5:46 buy products on the internet. 5:48 So it's supposed to do a series of tool calls 5:50 to actually identify what product to search, 5:53 and then have sequence of those actions 5:55 to actually complete the purchase for the user. 5:58 So each of these tasks require some notion 6:01 of grounding in the real world. 6:02 And without that grounding, it's not 6:04 possible for the model to keep making progress. 6:07 A second benefit of having the model 6:09 explicitly do reasoning and action 6:12 steps is that it will allow the model 6:15 to give traces that are more interpretable, 6:17 and interpretable allows humans to really trust 6:21 the model's responses. 6:22 It's almost like getting the model to break down 6:25 what it's doing step by step. 6:28 So what you might be familiar with is that, 6:30 if you prompt the LLM and you ask it to think step 6:32 by step, which we were showing in earlier examples, 6:36 then you will get some form of a reasoning 6:38 trace where it shows its intermediate steps. 6:40 And then, if you ask the model to, say, do a tool, 6:43 call go search, for example, Google or find the temperature, 6:48 it will go call an API, and it will 6:50 get a bunch of observations. 6:51 So that would be action only. 6:54 Typically, you will have a list of actions 6:55 if you have a list of tool calls that you are valid. 7:00 What ReAct does is it will take a goal 7:03 and have some form of reasoning, so that will be the left loop. 7:06 And then based on those reasoning chains, 7:08 it will go and act. 7:09 That will be the loop in the right. 7:14 So what this is setting up is a full agent in itself, 7:18 where you have some observations about the world 7:20 and then some set of actions based on what's the context. 7:23 And if you were to, for example, take this entire context 7:26 with observations and the set of actions that have happened 7:29 in the past, typically, you could go learn a policy, 7:32 but that would require a full loop, which 7:35 will be complex and expensive. 7:36 But you can actually start doing this in language space, which 7:39 is what ReAct is doing, where it's basically just generating 7:42 the thoughts in the language space, 7:45 and those thoughts are not affecting the environment. 7:47 And then based on those thoughts, 7:48 it's then going and taking an action, just 7:50 in the prompting space. 7:51 So it's a very simple abstraction 7:53 of how to prompt the models to start 7:56 thinking about these things and generate 7:59 actions, which are valid. 8:01 One interesting concept I'll mention here 8:03 is that, when you get the models to generate actions 8:06 in this particular slide, typically, 8:09 how do you know that these actions are valid? 8:11 So one typical way in which you get 8:13 these models to generate correct actions is to give them a more-- 8:20 to frame the problem more as a classification task. 8:22 So you say that this is the valid subset of actions, 8:25 and then use that as a way to say, OK, here is your reasoning. 8:28 Here is the state of the world, and here 8:29 is-- what are the valid set of actions to go take next? 8:32 And then based on that, it goes and decides, 8:34 OK, this is the action to take, and then 8:36 that generally provides a more grounded response. 8:39 And in fact, this was also used in another paper 8:42 that we worked on in collaboration at Google 8:44 with robotics folks, where the state of the world and all 8:48 of that requires very valid actions. 8:50 So you can't really operate without having 8:52 a list of actions. 8:53 So that provides one way to really grounded knowledge 8:56 and what actions to take next. 8:59 So in terms of implementing this idea in a very simple way, 9:05 this was done with-- 9:07 this is an older model, but with the PaLM model 9:09 with a frozen LLM. 9:11 In few-shot examples, the model was 9:13 given some actions, some thoughts, and some observations, 9:15 and then asked to go take the next action. 9:18 So for the reasoning traces, it was basically 9:20 seeing actions and thoughts. 9:22 I'll show you an example. 9:23 And for the decision tasks, it was basically looking at, 9:26 what's the state of the world? 9:28 What is the thinking? 9:29 And then deciding whether to reason 9:31 and whether to take action next. 9:33 So you can either go and reason, or you 9:34 can decide whether I have to reason or take action next. 9:38 So let's look through an example and walk through it carefully. 9:43 So aside from the Apple remote, what other device 9:48 can control the program Apple remote 9:51 was originally designed to interact with? 9:53 So this is basically a question from HotpotQA. 9:57 If you basically just do standard prompting, 9:59 the answer that will come back will not be correct. 10:02 Now, you can do Chain of Thought where you basically 10:04 ask the model to go think step by step. 10:06 It's showing you all the intermediate steps here. 10:08 However, the answer is still not coming out 10:11 to be completely correct. 10:15 Now, let's see. 10:16 If we ask it to take some actions where 10:18 it can do some grounded knowledge, 10:20 and it can do some search calls, so it goes and sees that, 10:23 OK, the thing that I'm really needing more information on 10:26 is Apple remote. 10:27 So it goes and does a action on Apple remote. 10:30 It gets some observation. 10:32 Then it goes and does another search call 10:35 on Front Row, which is some text that is there in observation 1 10:40 and then continues to follow on. 10:42 And then it realizes that, OK, this 10:45 is a discontinued something. 10:46 So as a result, there is not really a good answer 10:49 that it can give. 10:51 Now, how does thinking and action interleaving those traces 10:56 help? 10:57 So here, it thinks first that, OK, 10:59 I have to search for Apple remote, 11:00 and then find the program it was originally 11:02 designed to interact with. 11:04 So it first goes and searches Apple remote. 11:06 It finds the information that-- it's 11:08 a remote control introduced very early on, 2005, two decades ago. 11:13 And it was originally designed to control the Front Row media 11:16 center. 11:17 Then it thinks, OK, so I need to search Front Row next. 11:21 So it goes and searches for this. 11:23 It's not able to find it. 11:25 So then it comes up with a new terminology 11:27 called Front Row software. 11:28 And then it searches that, and it's 11:29 able to figure out, OK, it's a discontinued media software 11:32 center. 11:33 And then it thinks about it again, 11:36 so it's able to give an answer based on that. 11:39 So basically, the reasoning process 11:41 is an abstraction, which allows it to come back 11:43 with better answer than just taking actions 11:46 and then coming back with answers in between. 11:48 And interleaving the thought and action and the observation 11:52 aspect is basically allowing the model to behave better. 11:55 Now, what is happening in the LLMs 11:57 today where, for example, if you go in the thinking 12:00 mode of, say, Qwen or other models, 12:03 you will see this happening automatically. 12:05 So if you actually turn on the thinking mode 12:06 and say, Qwen models, you will see 12:08 all of this starting to happen because the models have already 12:12 been distilled on these traces. 12:13 So the thinking fusion stuff that 12:15 happens in several of the open-source models 12:17 already captures these things. 12:19 So they've already learned how to do, 12:20 for example, tool calls and so on. 12:23 Questions? 12:27 Yes. 12:28 How does the model know what it knows, what to search, 12:33 or when to search? 12:35 Because clearly, you don't want to search on 1 plus 1. 12:38 But how do you determine the search? 12:40 How do you determine-- 12:42 I mean, here, it's searching, and then 12:43 it's basically coming up with-- here, it's getting a feedback. 12:47 It doesn't know what it knows. 12:49 This is from interaction. 12:51 But you already have been trained on Front Row Wikipedia 12:55 page. 12:55 You shouldn't need to search. 12:57 Yes, It doesn't need to search. 12:58 Yes. 12:58 But does the model know whether it knows this? 13:01 So I think there's contradictory set 13:04 of opinions on whether the model knows what it knows. 13:07 There is one set of people that will tell you 13:09 that models are very confident in knowing what they know. 13:13 But I think the other side is that, typically, 13:15 if you were to ask the model to rate its output of, 13:18 are you confident about an output? 13:20 Typically, the models are overconfident. 13:23 They're not well-calibrated. 13:25 I mean that's a research problem that is still being solved. 13:28 So as someone who's designing a application, what 13:34 you are looking for is not so much 13:36 knowing whether the model knows what to know, 13:38 but getting the model to use the right set of tools 13:41 so that you have grounded knowledge. 13:43 Does that make sense? 13:44 Yes. 13:45 Yes. 13:46 With the ReAct on the right, when 13:49 you're saying with the left loop, 13:51 you elicit the reasoning to then append to original knowledge. 13:55 Will it produce the reasoning of the one or two 13:58 or three or four initially, and then it goes and just acts 14:03 for each one, or does the act influence the next reasoning 14:06 step? 14:07 It's interleaved, so it's thought one, then act one. 14:11 So it's happening in an interleaved fashion-- 14:13 thought one, act one, thought two, act two. 14:17 So it's basically each one is informing the next. 14:20 It's very much like the human process. 14:24 You're thirsty, so you, basically, 14:27 start walking towards the kitchen. 14:28 You get to the kitchen, maybe water is out. 14:30 So then you decide the next step based on that. 14:33 And so you think that maybe I have to go to the supermarket 14:35 to go do something. 14:36 So it's very sequential in that respect. 14:40 You can do parallel sampling and all of those tricks, 14:43 but then you have to explicitly design that process. 14:49 Yes. 14:49 There's no guarantees as a search result. [INAUDIBLE]. 14:52 So what do you [INAUDIBLE] observe [INAUDIBLE] 14:56 is the search result? 15:00 That's a great question. 15:01 If there's contradiction between the results, 15:04 then what do you want the-- 15:08 I think this is a question that the user has 15:10 to decide if there's not enough consensus. 15:15 So typically, one of the things that we covered in second class 15:18 was that we can get more confidence 15:21 by doing, say, majority voting. 15:24 So you basically say, OK, or you have 15:27 some way of validating the answer 15:28 before you give it to the user. 15:29 So for the end application, what you care about is 15:32 the answer is valid. 15:33 So you need more guardrails. 15:38 But hallucination with search results 15:41 can be controlled much better than just 15:43 depending on the internal state of the model. 15:47 Other questions? 15:49 Yes. 15:50 Maybe this is not answerable, but also [INAUDIBLE] from the-- 15:54 I wonder if you have looked at the KV-cache representation 15:57 of the thought versus the observation 16:01 and see if there is any sort of interleaving patterns 16:04 between them, given that the observations are generally 16:07 given by the environment. 16:08 They are not generated by the LLM itself. 16:10 So I wonder if there would be these changes in between. 16:15 That might be your research project for this class. 16:18 There are open-source LLMs that you can work with. 16:24 Yes. 16:25 It's just a little termination, like technology [INAUDIBLE], 16:29 but I know this is like the action space of ReAct 16:32 from where it's the actions plus thought source. 16:36 Why do we just think of thoughts internal 16:40 to the policy model, and maybe the action 16:43 is just like search or finish? 16:47 Why do we say source is also part of action space? 16:51 We're not calling it part of action space. 16:53 We are calling it a separate thing. 16:55 It's like actions are on the right loop, 16:58 and the reasoning traces are on the left loop. 17:01 Are you talking about this one? 17:06 So I think because the language models are trained on language, 17:10 typically, they benefit from having the reasoning tokens 17:12 in the right abstraction as a way to output the right action. 17:19 So that's the main reason to have those tokens. 17:21 I mean, if you had intermediate representations, 17:23 then maybe it doesn't matter. 17:28 So let's take a look at some results. 17:30 So these two tasks-- 17:31 HotpotQA is a multi-hop question-answering task 17:35 over Wikipedia, and then FEVER is more of a fact-checking task. 17:39 The action space in these two tasks 17:40 are very much focused on, say, search over Wikipedia pages 17:44 or looking up a certain string or finishing 17:47 that, OK, I finished whatever task I was given. 17:50 The action space is actually very simple. 17:51 It's basically simulating the human interaction with Wikipedia 17:56 here. 17:57 So the baselines here are simple, 17:59 where you're basically building a prompt. 18:01 The standard prompting that you will 18:03 see as simple prompting, no thoughts, no actions, 18:06 no observations. 18:07 Chain of thought where you're basically asking the model 18:09 to think step by step. 18:11 Chain of Thought, self-consistency-- 18:13 SC stands for self-consistency-- where 18:15 you're basically doing majority voting or self-consistency. 18:18 Both of those terminologies are exactly the same thing. 18:20 And then Act refers to there is no thinking in between. 18:23 You're basically just taking actions. 18:25 And there is two variants you will see in the table 18:28 where you can fall back to Chain of Thought, 18:31 self-consistency if ReAct fails after a certain steps, 18:34 and then this. 18:35 If the majority answer occurs less than half the time, 18:37 then it can back off to ReAct. 18:39 So you'll see this in the results. 18:41 So this is an example of when they prompted this large model. 18:46 What they found was that ReAct generally 18:49 does better than action only. 18:51 It did not always outperform Chain of Thought, 18:53 which was interesting. 18:55 It outperformed chain of thought in the case of FEVER. 18:58 But in case of HotpotQA, it did not always outperform. 19:01 But if you combined this Chain of Thought, 19:03 self-consistency with either fallback or the other way 19:07 around, then you do actually outperform 19:09 the standard techniques. 19:11 So what that roughly tells you is 19:12 that there is value in properly combining 19:14 the model internal knowledge with the external knowledge, 19:17 and then using reasoning as an intermediate way for the model 19:20 to arrive at the right answer. 19:22 So basically, that allows the model 19:24 to reason what is the right and accurate information 19:28 to retrieve here. 19:31 And it's basically showing that it's 19:32 able to use search as a tool to retrieve the right information. 19:38 And then another interesting aspect here is that-- 19:41 what this is saying is, when it's successful, 19:44 what is happening? 19:45 Did it get the correct reasoning traces? 19:46 And when it's failing, what errors are happening? 19:50 So one of the things that I want to highlight in this slide 19:52 without going through too much detail 19:54 is that, in Chain of Thought, hallucination 19:56 ends up being a major problem, a major failure mode. 19:59 But in ReAct, you basically end up having grounded information, 20:03 and the model is more trustworthy 20:05 because it has access to this external knowledge base. 20:07 So it allows it to make fewer errors 20:10 because it can successfully retrieve 20:12 the knowledge via search. 20:15 And then there is some difference 20:16 between prompting versus fine-tuning. 20:18 When you can fine-tune, then ReAct definitely does better. 20:21 And then, if you can do outer loop, 20:22 then it does even better, which we'll cover next week. 20:27 So then that was knowledge-intensive tasks. 20:30 If you were to look at decision-making tasks-- 20:32 and this can be, say, online web browser experiments 20:35 or, say, robotics experiments or gaming environments, 20:39 where we will look at WebShop as an example-- 20:43 you require the agent to purchase a product 20:45 based on user instructions. 20:46 For example, you're asking the model 20:48 to look for a nightstand with drawers, 20:51 and then you're basically evaluating 20:52 whether the model gets to success when it's interacting 20:55 in this environment. 20:56 It's a simulated environment. 20:58 And then you compare it as a baseline to imitation learning. 21:00 That's just fine-tuning. 21:02 And then imitation learning plus RL, 21:03 which is basically just supervised fine-tuning plus RL. 21:07 And the results that you see here 21:08 are impressive in the sense that, even compared 21:10 to, say, imitation learning or imitation learning 21:12 plus RL, having the reasoning abstraction with ReAct, 21:17 you get higher score, and you get higher success rates 21:19 compared to what's the baseline here. 21:22 So the score refers to immediate steps, 21:25 and the success rate refers to getting to the final-- 21:30 getting success on the action-- sorry, the task that the user 21:33 had assigned in this particular case. 21:35 So ReAct generally outperforms what 21:38 was possible with just imitation learning 21:41 or supervised fine-tuning before. 21:44 But it's still far below what human experts can do. 21:47 So there's still a lot of headroom compared to, say, 21:49 human experts, as you can see. 21:51 ReAct just gets 66.6 versus human experts 21:54 get 82.1 in this particular case. 21:57 And a lot of challenges happen with respect 21:58 to success rate, which is lower, because if you 22:01 get any of the steps, because this 22:03 is a multistep process, then the errors cascade over time. 22:10 So we covered ReAct first. 22:13 It's basically a simple method that 22:15 allows you to combine actions and thoughts. 22:18 But one of the challenges with this particular approach is 22:21 that, if you have very large action spaces, 22:23 then you need a lot more demonstrations that cannot fit 22:27 in context. 22:28 And then since you require a multiple reasoning steps, 22:31 your inference costs will be higher. 22:34 But overall, it's generally shown 22:36 to have better performance in question answering, 22:38 in fact-checking in decision tasks. 22:40 And then your hallucinations are lower, 22:42 and you get better decision traces. 22:46 So let's have a few minutes of discussion questions. 22:49 So what I will do is I will ask the class 22:51 to think through these questions and talk 22:54 to the person next to you to think through these questions, 22:57 and then we'll come back in two minutes. 23:03 Let's start with the last question. 23:05 Any takers for the last question? 23:09 By show of hands, any takers for the last 23:11 question to summarize what you guys discussed? 23:15 What happens if the environment is noisy, 23:17 if the feedback that you're getting is misleading? 23:27 You can make it up on the go if you haven't discussed this one. 23:34 OK. 23:35 So if the [? environment's ?] feedback noisy or incorrect, 23:38 you can add another layer of reflection. 23:41 So that's the agent to reflect on what 23:43 the environment gives you and then 23:45 reason about that [INAUDIBLE]. 23:48 OK. 23:50 Yeah. 23:50 I was thinking maybe it's important for the agent 23:52 to have an ability to backtrack whenever something doesn't make 23:55 sense, or if it's going through a chain of thought that 24:00 doesn't lead anywhere, just being able to backtrack. 24:04 Backtracking becomes important because the noisy feedback might 24:08 lead it down repetitive loops. 24:09 Yeah. 24:10 Or you could do external retrieval several times 24:13 and [INAUDIBLE]. 24:15 OK. 24:15 Yeah. 24:16 So the better confidence metrics. 24:18 So you need some way to build confidence over things. 24:22 You could have it perform the same task multiple times 24:26 and take the highest sample [INAUDIBLE]. 24:31 That makes sense. 24:32 Basically, you have a better understanding 24:35 of how confident the model is on this environment feedback. 24:39 OK, good. 24:41 Any ideas on the second one? 24:44 So ReAct is one way you can think 24:46 about reasoning and actions. 24:47 What other cognitive mechanisms should we 24:49 take into account if we are building these processes? 24:54 Yes. 24:54 I mean, humans, their permission is very different 24:57 depending on the task. 24:58 Sometimes we rely on our experiences in the past. 25:03 Sometimes we think through things step by step. 25:06 So it depends on the task at hand that we do that. 25:10 So just relying on ReAct and taking an action 25:13 doesn't necessarily lead us to the right place. 25:16 Sometimes you have to actually think and analyze 25:19 and do reasoning before we actually do any actions. 25:23 So I think what you're roughly saying 25:25 is that there is other thinking mechanisms that 25:27 perhaps are worth exploring. 25:29 One of them what I would phrase it as task decomposition. 25:32 Maybe the task is too complex, so maybe I 25:34 should decompose it into subtasks before attacking it 25:38 with these approaches. 25:40 Or maybe there is something to be 25:41 said about having parallel approaches to thinking, which 25:44 someone had brought up here as well, as to like, 25:46 is this approach better, or is this approach better? 25:48 And then having-- 25:49 Multiple memory, past experiences. 25:51 --have memory, past experiences. 25:53 So memory is another aspect. 25:56 OK. 25:57 Yes. 25:58 I think humans also try to fine-tune 26:02 what is the ideal ratio for reasoning to action. 26:05 For some tasks, you might learn over time 26:08 that it's better to think, give it much more reasoning time 26:11 before acting it out. 26:12 And it might be interesting to see if the paper actually 26:16 tries to train this on different tasks 26:20 to see what is the ideal outcome to compute 26:25 a ratio for each tasks. 26:28 That's a great point. 26:29 In fact, related to that, you'll find that some of the models 26:32 today overthink because they have really long thinking 26:36 traces, and yes. 26:39 And even for simple tasks, you'll 26:41 see them have very long thinking traces. 26:44 So there's definitely some interesting set of works there 26:46 of how to get the models to not think too much. 26:51 Let's go there. 26:52 Yeah. 26:53 Another interesting direction would 26:55 be, given the different model, the number of parameters 27:00 and maybe some model is better at being 27:03 able to do the whole tool calling or chain of reasoning 27:08 or even delegating tasks. 27:10 So being able to do based on what models you're 27:15 working with, supporting those different tasks. 27:18 Almost like building a compound system where you benchmark 27:21 subtasks on different models and then 27:23 delegate based on the strengths of the model. 27:25 OK. 27:27 That might be another project idea. 27:29 Cool. 27:30 OK. 27:31 Let's move to the second paper now. 27:34 So this one is called RLEF. 27:36 It's based on coding agents. 27:40 And what it's basically saying is 27:41 that, when you're building coding agents, 27:43 one of the things that you want to do 27:45 is give some feedback for the reinforcement learning loop. 27:47 And oftentimes, what really helps 27:49 is execution feedback, where we can execute and get 27:52 that feedback. 27:53 So RLEF was one of the first papers that 27:55 showed that with execution feedback, 27:57 you can actually get much better performance in coding LLMs. 28:02 Why do we care about this problem? 28:03 So how many of you use Cloud Code today? 28:07 OK. 28:08 That's almost majority of the class here. 28:10 So I don't think I need to explain 28:12 that a lot of the engineering and coding tasks 28:15 seem to have been delegated to these agents. 28:17 And the aspect that matters more in building these coding 28:21 agents to be extremely strong is that we 28:23 need to be able to understand the user intent. 28:26 And based on what the code was generated, 28:29 we need some feedback from the code 28:32 generated so that there is a way to iterate on top of that. 28:37 So what this paper shows is basically an end-to-end RL 28:41 fine-tuning framework. 28:42 The actions in this particular case are generated code, 28:45 and the observations that I was showing earlier 28:47 come from execution feedback. 28:49 Execution feedback is test feedback. 28:51 You're running some tests, and you're 28:52 collecting output of whether those tests pass or fail. 28:55 And you get binary reward based on whether those tests pass 28:58 or fail. 28:59 And then based on that feedback, you 29:01 can iteratively refine and fine-tuning time 29:04 and decide to make the model better based on that. 29:07 So PPO can be used for this approach. 29:11 The two techniques that this paper covers 29:13 are two-tier test strategy and then 29:16 hybrid token-turn level policy. 29:18 Both of those are very interesting. 29:20 So I'll cover the basic ideas here, 29:22 but I do encourage you to read this paper very, very carefully. 29:25 So the core of their framework is an iterative feedback loop. 29:28 They're using both training time and inference-time execution 29:31 feedback. 29:32 So let me explain the training time 29:36 and the inference-time feedback. 29:37 So at the very top in the purple block, what you're seeing 29:40 is that the model gets a natural language problem description. 29:43 For example, someone asked you to, say, write a Hello World 29:46 program. 29:47 Then it generates a code solution, 29:49 which then evaluates on some public set of tests. 29:52 If the code fails, the execution feedback is provided to the LLM 29:56 for another attempt, and this cycle 29:57 will continue until either the code will pass 30:01 or it will reach a limit of the number of turns it's 30:03 supposed to do this loop for. 30:06 And then based on this, whatever set of solutions 30:10 it has generated for passing solutions, 30:12 a private set of tests will determine what reward to give, 30:16 and then that will be used in the PPO training loop with RL. 30:22 So what here is happening is there 30:24 are two fundamental phases. 30:25 This is the exploitation phase of RL, where it's basically 30:28 exploiting the current policy through inference-time feedback 30:31 loops, like whatever policy model you have got. 30:34 And on the right side, you have update 30:36 of the policy based on the execution results, 30:38 and this allows continuous improvement for the model 30:41 to self-improve here. 30:44 So let's take a look at how this works. 30:48 You're given a simple example where 30:49 you're creating code to detect, say, palindrome substrings. 30:53 So that's what the problem statement on the top is. 30:57 In the first turn, it generates a basic solution, 31:00 but the implementation actually fails in the public tests, 31:03 and there's execution timeout, which is a very common issue. 31:08 After receiving this feedback, the model actually 31:12 generates a better solution where it does some optimization, 31:16 and now it does pass the public tests based on that feedback. 31:20 So you see that yellow block now submits the solution 31:24 to the private set of tests. 31:27 And based on that, it will go and try to improve the model now 31:30 that the solution is looking better. 31:33 One of the key aspects of their approach 31:36 was that, during iteration, the public test 31:38 will provide immediate feedback and will help it 31:41 do better solution development. 31:43 But this is a small subset for faster iteration, 31:47 and this is basically allowing the model to guide and get 31:51 better solution, while the private tests are completely 31:53 hidden during the generation process. 31:55 So basically, this separation allows 31:58 it to not go train on the public test, 32:02 but only use the private test as the way to get signal. 32:05 And this is one way in which the model cannot simply memorize 32:08 the test outputs, because it's getting the execution feedback. 32:12 So this, they found, is a very useful innovation 32:15 in the self-improvement loop. 32:17 A second innovation-- so I will actually not make you 32:20 go through the details of this RL algorithm. 32:23 This is basically PPO. 32:24 But the one thing I want to stress 32:26 is that when you're in a language model space, 32:28 you're outputting tokens. 32:29 So you're outputting one token at a time. 32:32 In the policy model, they are generating code tokens token 32:35 by token, so this gives you finer control 32:37 over the generation process. 32:39 But when they're using the value function, 32:40 they're computing this at turn level. 32:42 So this evaluation or the reward is 32:44 happening over the entire response, 32:46 and it's using the last token of the prompt. 32:49 So there's a single advantage value for all tokens. 32:51 So if you were to match this to where 32:54 we are in, say, RL algorithms, this 32:58 is closer to what GSPO would do in terms of giving rewards 33:01 to the entire sequence as opposed to doing it per token. 33:09 Yes. 33:10 I didn't understand how do you-- where do you 33:13 get the public license? 33:16 It's a test set. 33:18 You have a set of tests that you're keeping as public test, 33:21 and another set that you're keeping as private test. 33:25 Say your question again. 33:28 So on the left, you said that's part of difference 33:33 the user can give all sorts of program 33:38 descriptions [INAUDIBLE]. 33:41 Then this is for the outer loop. 33:43 So, for example, CodeContests, there 33:47 is a set of tests that are combined with CodeContests. 33:50 You have to design the outer loop 33:51 for the model to get better. 33:54 You can keep some subset of tests as public tests 33:56 on which you will do the inference-time feedback, 33:59 and then some subset of tests you can keep to train the model. 34:02 Oh, I see. 34:03 Thank you. 34:04 Yes. 34:05 What does it mean for the test to fail? 34:07 When the test is failing, that's basically 34:09 saying that the solution that the model generated 34:12 is not correct. 34:14 The model itself evaluates that, the correctness, or who 34:17 evaluates the correctness of-- 34:21 The test output is fed to the model. 34:24 You have failed, and then that's the execution feedback 34:28 that's going into the LLM. 34:29 You're solving this problem. 34:30 This is the output you generated, 34:32 and this is the feedback you got. 34:36 This is shown in the example that I've shown here. 34:39 This was the first thing, this code solution. 34:42 And then the public test failed, so you're appending all of this. 34:45 And then this is the execution feedback, 34:47 and then you're asking it to generate the second solution. 34:50 I was just wondering, what's the evaluation process 34:53 of a fail or pass? 34:55 That's an actual running of a test. 34:57 It's a Python test. 34:58 It's a simple Python function. 35:02 In your homework, you will actually-- not for code, 35:04 but for math, for any problems that you'll actually 35:09 try some of this out as not execution feedback, 35:12 but just identifying the error in the generated solution. 35:22 So why does this help or what's the benefit of doing this? 35:25 So one of the things that this paper showed 35:28 was that the y-axis is showing solve rates, and on the x-axis, 35:32 it's showing the sampling budget. 35:35 And this is for two different sets, for validation set 35:38 and for test set. 35:39 And even though the results are a little bit older 35:41 because they're on Llama 3.1, they clearly 35:44 showed that after RLEF and training on CodeContests, which 35:49 is competitive coding problems, you clearly 35:52 get much better results in terms of solve rate. 35:55 Basically, solve rate 10 at k means 35:57 that you're passing at least one out of-- you 36:02 have 10 solutions generated, and then you're 36:04 passing a certain number out of that. 36:06 So the solve rate definitely goes up with the RLEF. 36:09 And the other thing to note here is that this 36:11 is log scale on the x-axis. 36:13 So this is definitely improving with the execution feedback. 36:17 Now, why does it help is the other good question to ask. 36:20 So the base models typically don't benefit from access 36:24 to just faulty solutions and execution feedback. 36:27 What is really helping here is the fact that when 36:30 we are giving-- 36:33 when we are training on CodeContests, 36:35 we are showing it in each turn, what's the execution feedback? 36:38 And with that, you're basically also generalizing 36:41 to all the other benchmarks. 36:43 So it's the error in the solution, 36:47 that's what's helping the model start to get better. 36:51 And this becomes really clear in this particular example 36:54 where what they're looking at is number of errors in turn 1, 36:58 the number of errors in turn 2, the number of errors in turn 3, 37:01 and then the number of code changes being made 37:04 and the error type. 37:06 And this is for a smaller model, and this is for a larger model. 37:09 So what you basically find is that with RLEF, 37:15 as you go about iteratively in turn 1 and turn 2 and turn 3, 37:18 you have fewer wrong outputs. 37:21 And in subsequent turns, you basically 37:23 start to repair the output. 37:27 If you didn't have this iteration loop, 37:30 then you basically are not making edits that are correct. 37:34 So effectively, you are leveraging both the fact 37:36 that you have higher diversity in terms 37:38 of the number of samples that you're making, 37:39 but also your edits are more targeted 37:42 because you can look at where you made the error. 37:44 And that's one idea that you will use in your homework 37:47 problem as well. 37:50 So with the execution feedback, does 37:54 it also includes the faults you made, the errors? 37:59 Does it include any suggestions for how to fix? 38:02 It's just getting a sense of this 38:05 is the part that gave an error. 38:09 But these problems are very simple. 38:13 The [INAUDIBLE] problems are still, 38:14 I would say, not more than 100 lines of code, 38:20 so they have much smaller problems. 38:25 Yes. 38:26 How is the binary feedback enough to create the-- 38:32 I have a feeling that maybe these were very easy problems, 38:35 so that binary feedback was enough. 38:37 But for a harder problem, you would probably 38:38 also need to have the error trace or other metadata 38:42 to actually know how to debug it more efficiently. 38:45 That's quite possible. 38:47 I think that might make a very good project as well. 38:51 Yeah. 38:52 Especially in [INAUDIBLE], accuracy at the first time. 38:57 I feel like reward is binary reward for the final solution. 39:02 It means encouraging the model to just repair 39:08 and repair [INAUDIBLE] all together [INAUDIBLE]. 39:10 That's impossible [INAUDIBLE]. 39:15 What's the question there? 39:17 Oh, sorry. 39:18 So looking at the reward, is it just 39:20 binary reward or correct or incorrect? 39:23 But the model has a freedom to do [? that in model. ?] 39:27 How does this RL method encourage the model 39:30 to get it correct, get the problem correct, 39:32 in the first time? 39:34 So I think it really depends on the error it made. 39:37 And I agree with you that for more complex problems, 39:40 as was brought up, it is possible 39:41 that it will take multiple turns. 39:43 But if it was a simple enough problem, then-- 39:46 I mean, if you remember, there's inference-time feedback here 39:49 also. 39:49 And you're appending what is the error that it saw, 39:52 and you're allowing the public test 39:54 to pass before you send it to PPO. 39:58 So to some extent, it will have had a few chances 40:01 to go correct itself. 40:03 You can always increase the number of turns 40:07 to finish the problem, like any reward, you think that-- 40:10 Yeah, you can. 40:10 You can try all of those things. 40:12 I think these are-- 40:13 I mean, this is one abstraction of that problem 40:16 that I'm showing you. 40:17 And overall, I think what we want 40:18 to learn out of here is that the self-improvement loop works. 40:23 That's the first thing that we're learning, 40:25 and it can work in simple enough problems with binary reward. 40:28 And then potentially, I think, the question 40:30 is, between process reward models versus outcome 40:32 reward models, which works better? 40:34 I mean, what you're saying is that, 40:36 do we need to give feedback at every single step? 40:38 And it's possible that that is what matters. 40:40 Azalea covered it in last class, the outcome 40:42 versus process reward model. 40:44 So I think that debate is not completely solved 40:47 in terms of-- in each domain and each benchmark, 40:50 you might have to make different set of choices to help climate. 40:54 Yes. 40:56 I might have missed this, but why does 40:58 it encourage generalization? 41:01 Because I think-- [? isn't ?] context 41:03 is more difficult than human level and [INAUDIBLE]. 41:06 And that might be why it generalizes better. 41:10 And another is that, what about supervised fine-tuning, 41:16 would it be able to beat [INAUDIBLE]? 41:19 For example, [INAUDIBLE]. 41:20 Does supervised fine-tuning model or what always matters? 41:23 Supervised fine-tuning on reasoning traces 41:25 might actually still get you some 41:28 of the gains, as you probably have papers 41:31 out there saying that. 41:32 But I mean, this is still a question up for debate. 41:37 But I think with RL, you are able to solve 41:40 slightly newer problems. 41:41 So there is a little bit more generalization that's seen. 41:43 But with supervised fine-tuning, anything that's within domain 41:46 will definitely start to see value. 41:48 It's still the same loss function. 41:52 Do this have an ablation of the two-tiered [? task strategy? ?] 41:57 Because one way is you can just have all the tests in public 42:04 and then run this loop. 42:05 And then at the end, you do show. 42:07 I see. 42:08 So why did they do that? 42:09 What do you do with the [INAUDIBLE]? 42:11 On the final result after the [INAUDIBLE]. 42:17 I see. 42:18 I didn't see that ablation. 42:19 I think there must be, well, some amount of just 42:21 like they wanted to use more fine-grained feedback than just 42:25 the feedback of the public test so that there 42:27 is no leakage between the two sides. 42:29 Where would you care about leakage? 42:31 Why would you not care about the leakage? 42:33 I mean, they split the test into public, private themselves. 42:36 It's all test. 42:38 It's not in the training part of the model. 42:41 So why do you have to split it into two tiers [INAUDIBLE]? 42:44 I mean, in the outer loop, if you're training on the thing 42:46 that you're basically going to use as the feedback, 42:49 then that aspect matters. 42:54 We can take this one offline. 42:55 OK, sure. 42:56 I think there's a little bit of terminology gap around. 42:58 I see. 43:02 Yes. 43:03 In the ablation study, I was curious 43:05 why RLEF produces fewer wrong outputs, but more timeout 43:10 errors. 43:12 Oh, because it's probably-- 43:14 so it's pure wrong outputs because it's able to fix itself. 43:18 But then if it ran out of time in [INAUDIBLE], basically, 43:24 the tests ran out of time because the solution was still 43:26 not correct. 43:30 A lot of this is very domain-specific, also. 43:37 OK. 43:38 I'm actually going to skip this slide. 43:41 But I think the main thing that I want you folks to take away 43:43 is that there is a way to incorporate 43:46 execution feedback to improve code generation as a domain. 43:51 And competitive programming tasks 43:52 is one place where it has shown promise, 43:55 and it does generalize to other benchmarks in the code 43:57 generation realm. 44:02 So let's take a look at, say, question 2 44:06 and discuss that for one minute among your peers, 44:10 and then we'll come back together as a class. 44:15 Any takers for this problem? 44:18 The code base doesn't fit the context window. 44:23 So one thing I know is that some tools, what they'll do 44:26 is they'll do something similar to the previous paper we're 44:28 discussing, and they'll think about what information 44:32 they need. 44:32 They'll have access to some search tools, 44:34 and then they will continue iterating on 44:36 that until at some point, they decide 44:37 they have enough information. 44:39 And then they move on to something 44:40 like what we are discussing. 44:44 That definitely makes sense. 44:46 Any other takers? 44:53 Yes. 44:54 I think he would make this the code of each element [INAUDIBLE] 45:00 for summary so that [INAUDIBLE] be something [INAUDIBLE], 45:03 and then together, they can fit in the comments window. 45:10 Any other ideas? 45:12 Yes. 45:13 I don't know if this applies here, but in a regular setting, 45:17 I feel like you can build a graph-regularized version 45:20 of your code base and then do a similarity based 45:24 on what you can find them, and again, do 45:27 another iteration on that, just simple [INAUDIBLE]. 45:30 And then you can have in the loop what 45:34 you said about verifying and see if you have enough information. 45:37 And then, if not, you can do that, because in that case, 45:41 you can use the summarization. 45:42 You can use the things that you're extracting 45:44 and then come up with the answer. 45:49 So this is actually the kind of thing 45:51 that actually runs in Cloud Code. 45:53 It has to go search for what is relevant 45:56 and then actually apply the code, and it's there. 45:58 And one of the benchmarks that targets this is like SWE-bench. 46:02 So all the ideas that you folks have proposed 46:04 in the realm of searching for something, having some summary, 46:08 doing some representation, all of those 46:10 are perhaps the first step to solving that problem 46:13 before you now do the test, what code patch to apply, 46:18 and then whether it has passed or not. 46:21 So Code Monks paper from Azalea's lab 46:23 also targets very similar ideas here. 46:30 Let's move to the last topic. 46:32 The last topic is constitutional AI, 46:34 where we'll have the model learn from AI feedback. 46:37 And the thing that we are improving 46:39 about the model that is harmlessness, 46:42 or its ability to generate harmless outputs. 46:46 So we want it to be helpful, but we also want it to be harmless. 46:50 So what happened with LLMs when they are base models 46:53 is that you can give them feedback in a lot of ways. 46:57 A typical way to give feedback when 46:59 you are trying to build a chatbot 47:00 would be that you look at the outputs, 47:02 and you ask humans to rank them in, is the model output correct? 47:06 So you show it two different outputs. 47:08 You ask it whether it's correct, whether it's useful, 47:11 and whether it's specific enough. 47:12 And then based on those human preferences, 47:15 you can build a reward model, and then you 47:17 can use that to hillclimb. 47:18 So here is a Python fine-tuning with RLHF 47:22 versus just Python fine tuning. 47:24 And as you increase the model size, 47:26 you clearly see better responses with RLHF. 47:29 So basically, that's this whole notion of you 47:31 can use these preference responses 47:36 from humans as human feedback to hillclimb on the model. 47:40 Now, that doesn't scale very well, 47:42 because if you have to collect tens of thousands 47:44 of human labels, that's extremely 47:47 time consuming and tedious. 47:49 Imagine that you generate all these model outputs, 47:51 then you have to ask humans to rate them. 47:53 So one of the very interesting principles, 47:56 given that the models can reason, 47:57 was to use this notion of constitutional 48:00 that Anthropic came up with, which was very much based 48:02 on human-written principles. 48:04 So they basically used constitutional AI 48:06 or human-written principles called the Constitution, 48:09 and then used them to improve the model by describing 48:12 what is the desired behavior. 48:14 So humans don't need to be in the loop, 48:15 except to write that Constitution. 48:17 And the reason this notion of rules works 48:19 is because the models get better at instruction following. 48:22 So if you ask the model to format 48:25 the response in a certain way, the model 48:27 is able to follow that instruction. 48:28 Or if you ask the model that, does this response have 48:31 this particular behavior, then the model 48:34 is able to actually answer that in truthful ways. 48:37 That's what the constitutional AI 48:39 aspect will be able to exploit. 48:42 So what constitutional AI did was 48:44 it came up with a set of 16 principles, which 48:46 they call Constitution, and those will define the model 48:50 behavior, and they use that. 48:52 So these are going to be the set of prompts which 48:54 will elicit whether it's following the Constitution 48:56 or not, and then use that to critique the model itself. 49:00 I'll show you an example. 49:01 And then ask you to revise its response. 49:04 So you're basically saying whether-- 49:06 we're basically using some red-teaming prompts. 49:09 You're looking at the model output. 49:10 You're critiquing it. 49:11 You're getting a revision, and then you're 49:13 using that to fine-tune the model. 49:15 And so that's the supervised fine-tuning stage 49:18 for constitutional AI. 49:20 And then this is the RL stage where you're basically doing 49:24 the same set of responses. 49:25 And then using AI feedback to train a preference model, 49:28 and then using that to train the final model. 49:31 I'll show that in detail in just a second. 49:34 So before we go there, let's take a look 49:37 at what might be an example Constitution. 49:40 I will not read this word by word, 49:41 but I'll show you the basic examples. 49:43 So the critique request in this particular box diagram 49:48 that I showed you is that there is a red-teaming prompt. 49:50 The model has an output. 49:52 And then the critique request is basically 49:54 asking if the model output has something harmful or unethical 49:58 and so on. 49:59 And the revision request is basically 50:01 asking you to remove any of those responses. 50:04 A second critique request might be that, 50:06 does it have any gender bias? 50:08 And the argument about why that might be having a gender bias. 50:12 And then the revision request is, 50:14 can you remove any trace of that? 50:16 And the third one is more like, is it 50:18 inappropriate for young children? 50:20 And then what would it take for it to be appropriate? 50:24 And then the revision request is rewrite it. 50:26 So if you look at each of these, it 50:28 requires the model to be able to identify these behaviors. 50:31 And then the second thing it requires 50:33 is to be able to follow the instruction of now rewrite it 50:36 with this particular style or way, or removing this content. 50:43 And then the way this loop works is that humans come in and set 50:46 right a set of principles that will be used 50:51 for the self-improvement loop. 50:52 In the supervised learning stage, 50:54 you're basically fine-tuning the model 50:55 on this data generated by the self-critique and revision 50:58 stage. 50:59 So the critique request, as I showed you, 51:01 might be something harmful, something unethical, 51:04 and the revision might be, OK, remove anything 51:06 that's harmful and unethical. 51:10 And just with supervised learning with a number of turns, 51:14 you are able to-- 51:16 just with the number of revisions, 51:18 the harmlessness improves and the helpfulness 51:22 will decline if you just keep increasing 51:24 in supervised fine-tuning. 51:25 But overall, the helpfulness plus harmlessness 51:29 will improve monotonically, because if you're 51:32 trying to get the model to a certain set of fixed principles, 51:36 then it does become less helpful is what they're saying. 51:38 But overall, it's more helpful and more harmless 51:41 together is their claim. 51:44 And then in terms of the reinforcement learning stage, 51:46 first, they train a preference model 51:48 based on the responses in step 1 and the Constitution, 51:52 and then they fine-tune the LLM to maximize over this preference 51:56 model. 51:56 So they're basically getting the model 51:59 to respond in a way that's most thoughtful, respectful, 52:03 and cordial. 52:03 So this is basically the loop. 52:05 So there's a preference model that's being trained separately 52:07 with this Constitution, and then that's 52:09 what is being used to fine-tune the LLM to get better. 52:13 So instead of the RLHF loop where 52:16 you had a lot of human feedback, now you're 52:18 using this preference model, which is 52:19 trained with the Constitution. 52:21 Yes. 52:22 So all constitutions, maybe there may be some amendments. 52:26 If so, how do you make sure that you're cost efficiently updating 52:30 the constitutions for this model without having 52:33 to train all over again? 52:35 And if so, how do you make sure that the previous rules are 52:39 removed completely? 52:41 So there's two answers to your question. 52:43 So one answer is that [? AI ?] is generally done 52:47 as the last stage of model. 52:50 And typically, once you have a pretrained base model, 52:55 post-training has typically been a much smaller percentage 52:58 of compute, so maybe 5%, and that happens fairly frequently 53:02 for models to get updated. 53:04 So in that sense, if you do think 53:06 that the Constitution should be updated, 53:09 that happens at a certain frequency. 53:11 Overall, I think the question you're asking 53:13 is this notion of continual learning, 53:14 and how do we get the models to adhere to certain behavior. 53:17 I think that's an open research problem at this point in time-- 53:19 how do we get the models to forget certain set of knowledge 53:22 or to adhere to a new set of knowledge 53:24 or to follow a new set of rules? 53:26 So that's definitely something. 53:28 There are interesting answers to that problem where 53:30 you can do some canceling of getting 53:33 it to forget a certain set of knowledge 53:34 by looking at, say, interpretability methods 53:37 and so on of like, can you cancel that knowledge? 53:39 But then it's not proven that you can get 53:42 the models to forget something. 53:49 So in terms of how this scales, so this 53:52 is the number of training sequences they're using. 53:54 And on the x-axis, they're showing 53:56 the score of helpfulness Elo. 53:58 Elo means that humans prefer it in terms of helpfulness. 54:02 And then, similarly, on the right side, 54:04 they're showing harmlessness Elo where the humans prefer it 54:09 in terms of it being more harmless. 54:10 So now that we have replaced human preferences with AI 54:14 feedback, we are comparing it to, say, 54:17 just using helpfulness-based RLHF 54:19 or helpful-plus-harmlessness-based 54:21 RLHF, which is all human feedback-based, 54:24 and then just using constitutional AI 54:25 and constitutional AI with Chain of Thought. 54:28 And what they roughly show is that just getting the model 54:31 to evaluate its responses based on these principles, 54:34 you still get the model to be equally helpful, maybe 54:38 a little bit less helpful, but it's 54:40 much less harmless as a result. So the harmlessness scores 54:43 are much higher, and that was something 54:46 that the Claude models were extremely strong at. 54:51 And now that's a set of techniques that 54:53 are used across the models. 54:55 Yes. 54:56 The Chain of Thought has lower helpfulness Elo. 55:00 Say that again. 55:01 The Chain of Thought has lower helpfulness Elo. 55:06 Yes. 55:07 Results [INAUDIBLE]. 55:11 So it means that Chain of Thought actually worse, 55:13 the performance. 55:15 Possibly. 55:16 That's a good point. 55:18 Is there a measure? 55:20 And interestingly is that, on the right, [INAUDIBLE]. 55:23 So I think the correlation is less with Chain of Thought 55:26 and more with the fact that, as I showed you 55:28 in the last slide, this one, anytime 55:33 harmlessness is going up, there is 55:34 some amount of inverse relationship between these two. 55:37 So it's almost like-- 55:41 in this particular case, if it was more harmless, 55:43 then it does hurt the helpfulness. 55:46 But overall, you have to find the right balance 55:48 between these two. 55:50 Thank you. 55:51 Yes. 55:52 Let me go back to step 1 of the Constitution AI 55:55 to the supervised learning part. 55:57 So it said in the slides that we want 56:00 to fine-tune on the data generated 56:01 by self-critiquing relations. 56:03 So I want to understand whether this just simply means 56:06 that we want to just reinforce the log 56:08 likelihoods over literally the generations, the rollouts 56:13 of the L, of the original policy of L. 56:17 So basically, we just ask it to critique itself, 56:19 and then we just fine-tune this sampled rollouts so that they're 56:24 more likely. 56:24 That's what this stage is basically [INAUDIBLE]? 56:29 It's a stage that's used in a lot of other things 56:34 that you will see. 56:35 But what this is basically is doing a set of-- 56:38 these are traces that it's generated. 56:40 It's based on reasoning. 56:42 You want it to follow a certain set of principles, 56:44 and then it critiqued itself. 56:46 It revised itself, and it got fine-tuned 56:47 on this set of traces. 56:49 I mean, if you were training it for thinking, for example, 56:53 these would be the reasoning traces. 56:56 So I think being trained on-- it's being fine-tuned 56:58 on those reasoning traces, 57:01 But I guess it's interesting that there's not really 57:03 an explicit feedback-- 57:05 There is no feedback. 57:06 --because it's just asked to do this thing, 57:08 and then we just fine-tune it over what it has tried to do. 57:12 But as long as the fine-tuning is not very large 57:14 scale-- it's much smaller than what the base model was 57:17 initially trained on-- it will not 57:19 lose its initial capabilities, but the distribution of what 57:21 it will output, will start to get biased 57:24 towards this set of responses. 57:28 So I'm pretty sure the foundation [INAUDIBLE]. 57:34 The AI feedback stuff? 57:36 Yes. 57:36 How do we evaluate whether or not 57:38 they provide accurate feedback? 57:40 That's a good question. 57:41 So I think for the preference model, 57:43 when you do train this preference model, 57:45 you want to have some sort of test or validation set 57:48 to make sure that the preference model will do a great job. 57:52 So you do want to have some human. 57:54 But typically, what you'll do is you 57:55 will do a lot of AI feedback. 57:57 But for the preference model, you'll 57:58 also get it to give some scores on set of things 58:02 and then match it with humans. 58:03 So you do want to do some sort of consistency 58:06 with human for preference model. 58:08 [INAUDIBLE] 58:10 It's not [? quality. ?] 58:12 It's not scaled to 10,000 or that level. 58:15 You're basically able to do based 58:17 on such compositional principles. 58:26 So overall, in terms of results, what this roughly means is that, 58:30 if you were to plot the harmlessness Elo on y-axis 58:33 and the helpfulness Elo on x-axis, 58:36 the pretrained base model is basically the raw base model. 58:39 So it will have a certain set of scores, which are not very high. 58:44 But with standard RLHF, you will-- 58:47 if you were doing that with just helpful only, 58:49 you will start to push that frontier towards getting higher 58:52 scores on helpfulness. 58:53 And then, if you start to make it less harmless, 58:56 then that's what the orange curve shows you. 58:58 But then, if you do Constitutional 58:59 supervised learning, then you're still worse than doing RLHF. 59:02 But when you do Constitutional RL with Chain of Thought, 59:05 that's where you get the frontier in terms of the best 59:09 Pareto frontier in terms of tradeoff between harmlessness 59:11 and helpfulness. 59:12 And that's one of the key ideas of this paper as to you 59:16 can strike the right Pareto frontier 59:17 between helpfulness and harmlessness 59:19 in this particular case. 59:21 If you were to generalize these set of ideas across-- 59:23 I mean, generally anything you want 59:25 to do in instruction following where 59:27 you want the model to follow a subset of instructions 59:31 but they might be conflicting-- 59:32 then this style of principles applies. 59:38 And some follow-on work that went and built on top of that 59:42 includes the fact that there was another work that tried 59:45 to compare RLAIF with RLHF. 59:48 There was self-refine, where it talked 59:49 about iterative refinements with self-feedback. 59:52 One comment that's worth noting here 59:54 is that oftentimes getting the model 59:57 to critique itself can be harder, 59:59 so having a consensus of other models 1:00:01 to critique the model sometimes works better 1:00:04 because the models might be overconfident 1:00:06 and not knowing what they know. 1:00:09 And then there was another paper that 1:00:11 tried to teach models to self-correct 1:00:13 via reinforcement learning. 1:00:14 So this is an active area of research, 1:00:16 just trying to get the models to be better at self-correction 1:00:19 just beyond just the helpfulness and harmlessness paradigm. 1:00:23 So that might be another set of project ideas 1:00:25 that you might want to explore. 1:00:28 So let's recap and then we'll close out the class. 1:00:33 So in terms of the ReAct approach, 1:00:36 what we first looked at was this notion of, 1:00:39 when we talk to LLMs, we can get them to give model outputs 1:00:42 and what we wanted these model outputs 1:00:44 to be preferable by humans, which has been a focus. 1:00:48 But what ReAct starts to look at is this notion of, 1:00:51 can we get the models to act in real world? 1:00:54 And based on that feedback and combining 1:00:57 that with reasoning, we can get them 1:00:58 to be grounded so they can give better answers to questions, 1:01:02 can check facts better. 1:01:03 Or if we put these LLMs as decision makers in, say, gaming 1:01:08 environments or in other environments related to, say, 1:01:11 web browser agents, and so on, then they 1:01:14 can make useful decisions. 1:01:16 And one interesting aspect of doing it in ReAct style, 1:01:19 in language space, ends up being that the decision traces are 1:01:22 extremely interpretable. 1:01:24 There are several works on top of that 1:01:26 try to combine this reasoning and acting in different ways 1:01:29 and trying to get the models to do better tool calling 1:01:32 to the point where this tool calling is 1:01:35 innate in several models that are out of the box today. 1:01:39 But this is the building block of how 1:01:41 to get the language models to combine reasoning and acting. 1:01:44 A second aspect that we covered today was this notion of RLEF. 1:01:48 So in coding agents, getting the models to act in the correct way 1:01:54 requires some form of execution feedback, 1:01:56 and you want to iteratively incorporate 1:01:58 that feedback so that you can generate the right code 1:02:02 solutions. 1:02:03 For code, one of the best ways to get execution feedback 1:02:05 ends up being tests, unit tests. 1:02:08 So that's one way the models have been used to self-improve, 1:02:13 and this also reduces the amount of budget you need to get 1:02:17 the models to get better and get state-of-the-art performance 1:02:21 in, say, CodeContests, and other competitive programming tasks. 1:02:24 So that's one aspect of just getting execution feedback 1:02:29 in the coding domain. 1:02:30 And then finally, what we covered 1:02:31 was this notion of constitutional AI 1:02:33 where we are getting the model to follow 1:02:35 a set of constitutions, which is a set of principles written 1:02:37 by humans, to generate feedback for self-improvement. 1:02:41 In general, you can also generalize it to anything 1:02:44 where you can get the model to follow a set of rules. 1:02:46 And then based on whether it's following the rules well or not, 1:02:49 you can get a self-improvement loop going. 1:02:51 So hopefully, with these set of papers, 1:02:54 you got a sense of how we can build a feedback 1:02:57 loop on top of the models. 1:02:59 And when the feedback loop has enough signal, 1:03:01 then you have a way to improve the model beyond just what is 1:03:06 the data it was trained on. 1:03:09 And these were several examples of where this 1:03:11 has been shown to work well. 1:03:13 Sorry, this slide was not meant to be there. 1:03:16 So that would be the end of the lecture, 1:03:19 but let's take any questions. 1:03:30 OK. 1:03:30 This is a really naive question, but are there-- 1:03:35 so because a lot of human cognitive thinking 1:03:38 is very paralleled in how we're designing elements, 1:03:41 is there actual, very systemic research 1:03:45 to combine cognitive behavioral science with elements? 1:03:49 Or how do we think of new procedures 1:03:53 to innovate elements like this? 1:03:56 That's a great question. 1:03:59 I think where we are right now, that's what we are-- 1:04:03 we're using parallels from how humans solve problems 1:04:06 because we are in this space of reasoning LLMs. 1:04:10 We're in the era. 1:04:11 But I think at the end of the day, 1:04:14 we're still trying to solve the problem of what 1:04:16 would get the models to hillclimb without explicitly 1:04:21 saying in words what's the reasoning space. 1:04:23 So I think the closed-source LLMs have done better 1:04:29 than just saying that this is their reasoning space in words. 1:04:35 So effectively, what I'm trying to say 1:04:37 here is that you can do a lot of techniques 1:04:40 where you can be like, I need to break down the task step 1:04:43 by step, which is something that Azalea 1:04:46 presented in the first lecture. 1:04:47 I need to decompose the task. 1:04:49 I need to analyze the task, and send it in parallel loops. 1:04:52 And that's how humans solve problems. 1:04:56 But at the end of the day, you also want 1:04:58 to answer the question of, can that whole process 1:05:01 of search in the reasoning space be automated, 1:05:05 and is there a well-defined process to do that search? 1:05:09 And if the search space is well-defined, 1:05:12 then you can automate it. 1:05:14 The challenge of why we're doing it, not in an automated way, 1:05:17 but in this language spaces, because a lot of tasks 1:05:21 that we're giving to these models 1:05:23 is not in a well-defined search space. 1:05:25 A good example of that would be that, 1:05:27 if you were basically in gaming space, 1:05:29 where the action space is well-limited 1:05:30 and the search space is well-limited, 1:05:32 then you can define models to go figure 1:05:36 out what the search space is and how 1:05:38 to explore that search space. 1:05:39 But if we are talking about experience 1:05:42 and going and experiencing actual environments, 1:05:44 whether that's search tools or everything else, 1:05:47 you have to collect observations there and then use that. 1:05:51 So that's why you're seeing this. 1:05:57 So far, we've been talking about improving LLMs 1:06:00 through different techniques. 1:06:01 I wonder how those techniques are transferable to improving 1:06:04 agents. 1:06:06 You will see that in some of the subsequent lectures. 1:06:09 But I mean, all of this-- 1:06:12 what's your definition of agents versus LLMs? 1:06:14 Oh, is basically some LLM that are [INAUDIBLE] 1:06:19 and memories and sketches that can carry all of us. 1:06:24 So I think this is starting to represent 1:06:26 what I would call an agent, but not with-- doesn't have 1:06:29 sessions, doesn't have memory, because one 1:06:32 of the fundamental building blocks of agents 1:06:34 is, how do we get the LLMs to respond in a certain way? 1:06:37 So this ends up being very useful to understand. 1:06:43 So for frameworks, like ReAct, my understanding 1:06:46 is that they have some intuition of what 1:06:48 kind of [INAUDIBLE] thing calls a reasonable help the model. 1:06:53 But is there a relationship that only [INAUDIBLE] 1:06:56 a certain domains, and then does it 1:06:58 help versus [INAUDIBLE] the other [INAUDIBLE]? 1:07:01 So now that we have the more RL-based post-training methods, 1:07:06 will this handcrafted framework become obsolete? 1:07:12 Yes and no. 1:07:14 The reason the answer is yes is because, 1:07:17 if you can define the search space of what 1:07:20 to go explore, then yes. 1:07:22 No, because defining the search space 1:07:24 for all tasks of what to go explore is not well-defined. 1:07:29 If you wanted to do the task of an accountant, 1:07:32 if you wanted to build a finance agent, 1:07:34 if you want to build a legal agent, 1:07:36 what is the set of steps they follow? 1:07:39 That's very domain-specific. 1:07:42 But it's just generating tokens. 1:07:44 [INAUDIBLE] is also just tokens. 1:07:47 So you're just generating rewards 1:07:49 using the reward to supervise them. 1:07:51 So I think in an abstraction, I can answer it as yes. 1:07:55 But the specifics still matter. 1:07:57 So the reason ReAct as an approach matters 1:08:00 is because they are basically saying, 1:08:01 OK, I can reason that I can take action, then I can go reason, 1:08:04 I can take action. 1:08:05 So it's basically defining the sequence and the workflow. 1:08:07 If you remember from the first lecture 1:08:10 when we defined the agent, we said 1:08:11 it's basically some abstraction of how 1:08:13 I would define that workflow. 1:08:15 So how I would generalize the whole ReAct approach 1:08:18 to an agent is I would say, OK, this 1:08:19 is the workflow that I need to follow 1:08:21 to accomplish the task that the user has given. 1:08:24 But the workflow to accomplish the task 1:08:26 is still very domain-specific, so that's 1:08:29 why it's very hard to generalize across domains for everything. 1:08:34 Yes. 1:08:35 Well, I have a question about the formalism on slide number 1:08:38 12, the methodology setup page. 1:08:41 One question I have-- 1:08:44 is that there, the way that you define 1:08:47 C subscript t, the context, is the context 1:08:50 just simply being treated as a state? 1:08:52 Yes. 1:08:53 OK. 1:08:54 I mean, the reason it's called context is that-- 1:08:56 I mean, RL state versus prompting is literally 1:09:00 what's in the context. 1:09:01 No, I totally agree with that. 1:09:03 And then the other thing is with ReAct, 1:09:05 because ReAct has this thought, and then there's this action. 1:09:08 But it seems like in the formalism 1:09:10 in page 12, the action, I suppose 1:09:14 that action would include, presumably, 1:09:16 both the thought and action because since the action 1:09:20 from that thought would just be extracted 1:09:23 from the output of the RL. 1:09:25 So in that particular case, it's being given the choice, 1:09:30 so it's basically either outputting a reasoning trace, 1:09:32 or it's outputting an action, and it's 1:09:35 making a decision whether to output one of those two. 1:09:39 So it's-- 1:09:41 Oh, I see. 1:09:41 So the action can be just like a standalone reasoning trace. 1:09:45 And then it stops, and then it does another action. 1:09:51 Yes. 1:09:53 This might be a naive question, but I was just 1:09:55 wondering what this harmlessness have certain features 1:09:59 of certain methods structures. 1:10:01 So we have [INAUDIBLE] obvious considerations 1:10:04 instead of just trying to saw in the feedback 1:10:06 to try to fine-tune the model to somehow improve 1:10:11 the harmlessness. 1:10:12 For example, what can we do [INAUDIBLE] instead 1:10:15 is [INAUDIBLE] can use [INAUDIBLE] 1:10:17 to generate certain methods to restrict [INAUDIBLE]. 1:10:25 Yes. 1:10:26 I think that is done to some extent. 1:10:30 At the same time, these models are trained on the internet, 1:10:34 so no amount of restriction is really going to get all of that 1:10:38 out. 1:10:41 So it's a game of when you are training on a lot of data, 1:10:48 that is the tradeoff. 1:10:52 The data is definitely filtering off data that happens. 1:10:58 If there are no further questions, 1:11:04 we'll call it a close of the class. 1:11:06 Thanks, everyone.