Title: Part 3 Robust Verification Authors/Date: Akanksha Bhardwaj, Azalia Mirhoseini — Stanford CS329A Video ID: p7TdPUcPoik | URL: https://www.youtube.com/watch?v=p7TdPUcPoik Playlist: https://www.youtube.com/playlist?list=PLangBM27OtEA ------------------------------------------------------------------------ 0:05 So lecture three is about verification. 0:09 In the last lecture when we talked about inference time 0:14 scaling and the generation verification gap, 0:20 basically what we discussed was that language model seems 0:24 to be able to seems to know the answer to many 0:29 of the hard questions, and especially with methods 0:32 such as repeated sampling, or other scaling test time 0:36 techniques, they can generate one. 0:38 But a question is, how do they. 0:41 How can we automatically select which answer is correct 0:45 or guide the model throughout the process 0:47 of answer generation. 0:49 So that leads us to verification where 0:52 which we are learning about today. 0:55 So in today's lecture, we are going 0:58 to learn about the following four papers. 1:02 And there is a progression of how the way 1:07 we approach verification changed or progressed 1:10 throughout the years. 1:12 So let's first start with training 1:15 verifiers to solve math problems. 1:18 This is a paper from OpenAI in 2021. 1:23 And the motivation for this fork was 1:26 that LLMs hallucinate and can confidently present 1:32 wrong solutions to the users. 1:35 This still is true to this day. 1:37 Four years later, of course, the model 1:39 have become significantly better. 1:42 There is also at the time, one of the contributions 1:45 of this paper was to introduce a new reasoning benchmark 1:50 around math. 1:51 And you can see here in blue one of these problems. 1:55 This is the type of problem that at the time it was hard for LLMs 2:04 to solve. 2:05 And this led to this new dataset. 2:08 And that has become kind of a big part 2:17 of many of the benchmarking tasks for LLMs. 2:20 And even to this day, this dataset is called GSM 8-K. 2:24 And even to this date, for smaller models and benchmarking, 2:28 it's still a very useful language modeling evaluation. 2:32 And you can consider it for the type of evals 2:35 that you want to define on your projects as well. 2:38 So it basically consists of 8,500 grade school math 2:43 problems. 2:45 And the goal was for this was to have the quality 2:49 and diversity of the problems. 2:51 But also importantly there was a focus on multi-step reasoning, 2:56 meaning that to solve the problems, 2:58 even though the problems were simple. 3:00 But the model required a few steps of reasoning 3:04 to get to the answer. 3:07 And you could also collect the solution in natural language as 3:13 opposed to pure math. 3:18 But the other contribution of this paper, 3:20 which we are going to spend more time on, 3:23 was to train a verification model 3:26 where the verifier outputs the probability that a solution is 3:31 correct. 3:32 So let's see about how this can be useful. 3:35 For example, the model that as humans we solve problems. 3:40 It would be very, very good for us 3:42 to have access to a rubric or a way 3:45 for us to tell whether we are correct or not. 3:48 And we kind of want to provide the same thing 3:50 through this verifier for the language model. 3:55 So now how these the verifier was trained in this case 3:59 is that assume that we have a question and a solution for it. 4:04 Then we do have a label whether is a binary label, whether it's 4:09 correct or incorrect. 4:11 And during the training process, the generator, the language 4:15 model, which is in charge of answering questions 4:19 here produces PSI, which is the solution to question QI, 4:25 and then we can generate-- 4:27 do repeated sampling and generate multiple, 4:29 in this case, 100 solutions per problem. 4:32 And then we create these labels based 4:34 on whether they were correct or not. 4:37 Remember that these questions and answers were selectively 4:41 designed by a lot of manual human supervision. 4:45 So we know what the answers-- the final answers 4:48 to these problems are. 4:49 So we can just, based on that, create these labels 4:52 for these problems. 4:54 And then given that we can train a verifier, 4:57 because now, we have a question, we have a solution, 5:01 and then we have a label for it. 5:07 Now, the way we use it at test time 5:09 is that at test time, again, we generate a lot of answers. 5:13 And then we use the verifier and look 5:15 at the score, basically, the log probability that 5:18 is in this model now, to see which one has a higher score 5:22 and use that for as our-- 5:24 and show that as our final answer. 5:31 So in this paper, the verifier was trained not just based 5:34 on the correctness prediction, but there was also 5:37 a language modeling objective. 5:41 So the language modeling objective 5:42 is also given, for example here, given these tokens corresponding 5:48 to the questions, the moment we get to the solution token, 5:53 we also add a loss that kind of urges-- like, 5:56 motivates the model to predict or reduce 5:59 the solution distance from the true labels for that solution. 6:08 So in this paper, they said that it 6:10 helped them to have two losses. 6:12 One is the binary loss and the other one 6:14 was the normal language modeling, 6:17 next prediction token loss. 6:20 And the verifier itself, the architecture of the verifier 6:23 is also a language model. 6:26 But also, this language model has a small scalar head 6:30 that outputs the binary prediction on a per-token basis. 6:38 So basically, on a per token we are 6:40 saying how close was the solution to the true labels 6:46 generated by humans. 6:50 And then the tokens in questions are masked out 6:53 because we are not optimizing for that. 6:56 So we generate the question. 6:58 We pass in, fill in the model with the questions, 7:01 and we ask it to generate the solutions. 7:03 And we work with these two losses. 7:07 The way they trained the model was 7:09 to fine tune the generator language model for two epochs, 7:14 and then sample 100 completions per question, 7:19 label it whether they were correct or not, 7:22 and then train a verifier for a single epoch on this dataset. 7:27 So first, they did fine tune the language model on that dataset 7:32 on a portion of that dataset. 7:35 And this step maybe it's like debatable whether we still 7:40 need this step or not. 7:41 Because language models are already 7:43 very good at instruction following and understanding 7:46 math. 7:47 So maybe and for a lot of newer verifiers, 7:49 we directly go to the training the verifier objective 7:53 without any supervised fine tuning first. 7:58 So they did a number of ablation studies. 8:02 One was to create this label of correct or incorrect 8:08 at the sentence level. 8:12 So we have a loss that only concerns at each sentence level. 8:17 For example, after each period looks into the previous sentence 8:21 and looks whether this is a correct step given 8:25 of the correct step for the solution or not, 8:30 or they could do at a token level. 8:32 So this is very, very noisy kind of label. 8:35 Like basically we want every token to follow. 8:41 The pattern that leads to a final correct solution 8:45 at the end of this generation. 8:49 And then the way this is used at test time 8:52 is the final solution score is the score after the last token. 8:57 Because if you have a token label prediction per token, 9:03 at the end of the day, we want only one prediction 9:06 for the entire generation, whether it's correct or not. 9:09 And it turns out, in this case, they 9:11 looked it after the last token. 9:13 All right. 9:14 So we were talking about if you have a token, a label per token, 9:18 at the end of the day, we want one token, yes or no, 9:22 whether this answer was correct or not. 9:24 And this is done in this case by just 9:28 looking at the very last token. 9:30 And if that was a correct, then the entire answer 9:33 is assumed correct, and otherwise, it's 9:36 assumed incorrect. 9:37 So we predict for S1, S2, SM, we predict these labels. 9:42 And then we only look at the very last label 9:44 to generate and allocate a label for the entire solution. 9:50 So here is-- looking back at these problems, 9:56 when we get to the solution, we are 9:58 looking at the token-level prediction 10:00 by the train verifier. 10:02 And if it's green, meaning the verifier has a high score 10:06 and the red ones means a lower score in this case. 10:09 And in this case, there were some lower scores in the middle, 10:13 but towards the end, it became greener and greener, 10:16 and this was a correct prediction at the end. 10:20 At the end, the very last stage is a green one 10:23 and we assume the actual score and the verifier predictions 10:26 are in this case are matching. 10:30 In this other case, the last verifier is red. 10:33 It starts off well, but it has a. 10:36 There is something a mistake is made, 10:45 which is towards the end of this generation step 10:49 and the actual score is negative and the verifier prediction 10:52 is also negative. 10:54 So this is how basically this verifier is trained for GSM-8K 10:58 And then they did some ablations or comparisons 11:03 with a baseline where we just directly just fine tune 11:06 on this dataset fine tune meaning with supervised fine 11:10 tuning. 11:11 So we just take the base model and the next token prediction, 11:15 or we do this verification stage where 11:19 we generate 100 samples at each generation, 11:24 and use the highest score of the verifier 11:26 as the true as the delivered output by our system. 11:30 So the idea, what we can see here is that for both 6B 11:37 and 175b models, this one is a GPT-3 model, I believe. 11:44 For both of these, verification seems to work better, 11:48 especially as the training set size for the verifier increases. 11:53 So the verifier model eventually outperform 11:57 the fine-tuning-only approach. 12:04 And vice versa, if you have smaller data sets-- 12:07 in this case, less than 1,000 for the 175B-- 12:12 the verification is not helping much. 12:19 [INAUDIBLE] 12:24 Sorry, can you speak up? 12:31 I mean, the test summary of the 6B verification is still much 12:37 higher than the [INAUDIBLE] on the smaller [INAUDIBLE]. 12:44 The orange lines still-- 12:48 Isn't it lower-- 12:49 OK, I see. 12:50 --than 175B? 12:52 [OVERLAPPING SPEECH] 12:53 OK, let me go here. 12:56 So this is another result which is very interesting. 13:01 And this is something that could potentially 13:03 be a research project that you guys can look into. 13:07 So what they did is that they tested two approaches, 13:11 one having a larger generator and smaller verifier, 13:15 and two having a smaller generator and larger verifier. 13:18 And what they saw was that the larger generator helps. 13:24 The smaller verifier does better than vice versa. 13:27 And to some extent, this is intuitive, 13:29 because if we assume generation is naturally, or on average, 13:37 is a harder task than verification, 13:40 then this might make sense. 13:42 But finding the Pareto optimal of how these two can 13:47 be the sizes, with respect to each other, 13:50 would be a very interesting research question, especially 13:55 right now, as the base-model generators have improved a lot. 14:00 And also, we have access to a lot of verifiers. 14:02 On Hugging Face, there's a leaderboard 14:05 of verification, verifier models, or reward models 14:08 that can be tested against. 14:11 So again, this was the main message 14:15 from this ablation study. 14:19 Another result about test-time scaling that they're showing 14:26 is that, as they increase the number of completions 14:29 per problem, and then they use the verifier, 14:32 it seems like around 400 or so. 14:37 That's where they get a lot of benefit. 14:39 But after that, they're not getting a lot of benefit. 14:42 So if they sample 800 solutions and they 14:46 use the verifier against that, the verifier 14:49 fails to track what's best and what's 14:52 not, compared to the true distribution of the solutions. 14:57 So one question for you is that, do you 15:01 remember last time, when we talked about majority voting? 15:05 What was the range that we were getting out of that? 15:08 We showed that majority voting, although, is 15:11 useful to some extent as we increase 15:13 the number of samples per query, it fails, after some point, 15:19 to track. 15:20 And that's the generation-verification gap. 15:23 Does any does anyone remember that? 15:28 OK, so that was-- we saw that the failure happens after 10 15:34 to-- 15:36 we never got to 50 samples. 15:38 Even though the coverage goes up, the pass at one-- 15:42 or basically, how we can generate, 15:44 select the correct response-- 15:47 fails around 50 or less for the majority voting approach. 15:53 But here, they're showing that, up until 400, 15:56 they can increase the accuracy, and the verifier is still 15:59 useful. 16:06 Yes, but when they have-- so they 16:10 do have the true labels per problem. 16:13 And then, when they are testing this verifier, 16:15 the verifier ranks these at each point. 16:19 Say, for 400, these 400 solutions, 16:22 they take the highest score, and they 16:25 compare that against ground truth, 16:27 and that's what they're reporting here-- whether that 16:29 was correct or not. 16:31 My question is that, in the training time, 16:33 you said there's [? over ?] 100 [? fusion. ?] 16:36 At test time-- and training time as well, for the verifier, 16:40 but not for the generator. 16:42 OK. 16:42 For the training times, all are human [INAUDIBLE] 16:46 but the training-- 16:48 So, for each problem, the human has once generated an answer. 16:52 That's offline. 16:53 That's done once. 16:55 Then you take the model, generate 100 solutions 16:59 per problem, and you compare each of the solutions 17:03 against the ground truth, which is 17:05 generated only once by humans. 17:07 And you see whether they match or not. 17:09 That's how you create the label. 17:13 OK. 17:16 OK, any questions from this paper? 17:19 Yes? 17:20 I have a question. 17:22 After 400, it made sense. 17:25 The verifier is increasing results. 17:27 But after that, why does it drop? 17:29 Intuitively, why doesn't it stay constant? 17:33 OK, so the question is, why does it drop after 400? 17:37 It shouldn't be constant, because we are not 17:39 showing the coverage. 17:41 We are showing the output of this entire system, where 17:44 the verifier ranks these, say, at this point, 800 17:48 different solutions, and we are getting 17:50 the max score of the verifier. 17:52 But the precision of the verifier drops. 17:56 If two solutions are too close to each other, 17:58 but one is correct and incorrect across 800, 18:02 it can't distinguish between the two 18:04 as well as it could do for 400. 18:07 So that's where we get the drop in accuracy. 18:11 And in reality, when they're outputting-- 18:14 they created this system. 18:15 They said that they stopped at 100 samples 18:19 because that's where they get most of the gains anyway, 18:23 and that's where they're consistent. 18:26 [INAUDIBLE] it's getting longer and longer. 18:32 Does this mean we have more answers? 18:34 Repeated samples, so parallel. 18:38 So, same question. 18:40 You ask the model 100 different times to solve that. 18:47 Yes? 18:48 When they generate a label and these sort of sentences result, 18:52 how do you distinguish? 18:54 One could just be a perfect [INAUDIBLE] matched one-by-one-- 19:01 Because this is trained on a larger data set, 19:05 and because the verifier-- and they 19:07 consider both sentence-based and token-based. 19:10 It still works because the data that the trainer-- first of all, 19:15 the verifier is a language model itself, so it knows 19:18 a lot about this space itself. 19:20 But also, through this label training, 19:24 it has been trained to be able to distinguish 19:27 between correct and incorrect. 19:31 Yes? 19:33 Does it make sense to do test-time scaling 19:35 on the verification itself? 19:37 So the question is-- let me repeat the question because one 19:40 of you asked this before. 19:42 So the question is, does it make sense to scale the verifier, 19:46 do test-time scaling on the verifier? 19:48 Yes, and you're going to see that in the future slides. 19:56 Yes? 19:57 The plot that you showed verification versus 20:00 fine-tuning-- you showed that before, around 1,000 steps, 20:04 1,000 training size, right? 20:06 The fine-tuning actually does better than the verifier. 20:10 Does that happen again if you scale it to a higher training-- 20:15 Do you mean the fine-tuning versus verification? 20:19 Do they cross each other-- 20:20 Again, yeah, [INAUDIBLE]. 20:28 If you were to actually-- 20:30 if you do need that much data, for example, for a task, 20:36 do you ever-- 20:37 I mean, I think that you would. 20:39 If you have a ton of data, a whole lot of data, then 20:45 fine-tuning and verification-- probably, 20:48 they're going to be the same. 20:49 They'll converge to the same thing. 20:52 But the beauty of a verifier is that you are not fine-tuning 20:56 your base language model. 20:58 You're not making it really custom to one data set or one 21:02 specific task. 21:03 So that base model, the generator model, 21:06 remains general. 21:08 And then, now, you have a verifier 21:09 that you can lead the base model. 21:18 Now let's go through the second paper, another paper by OpenAI. 21:25 Let's verify, step by step. 21:32 And so this came about a couple years, less than two years, 21:38 after that first paper. 21:40 Again, LLMs have had the same challenges. 21:43 They hallucinated. 21:46 If they're solving a problem step by step, 21:49 and they make a step in a single-- 21:51 a misstep early on-- it can derail the entire answer, right? 21:56 So the solution that they were proposing here 21:59 is that they consider these two kind of approaches 22:04 to reward modeling. 22:05 One is outcome-based reward model, 22:09 and the other is process-based reward model. 22:12 So the outcome-based one, like we saw earlier, 22:15 in the previous paper-- it generates a reward 22:18 to the entire solution. 22:20 It's about the correctness of the solution, 22:24 of the entire solution, whereas the process one assigns 22:28 a different reward per step of the optimization problem. 22:36 And so let's take a look exactly how this training can 22:40 be done for the ORM versus PRM. 22:42 So we have a math problem. 22:44 The generator-- instead of generating 100 solutions, 22:48 right now, we are only looking into one solution. 22:52 And it has a bunch of steps-- step 1, 22:54 2, and all the way until we get to the final answer. 22:59 And then we have the ground truth, 23:01 so we can know whether the final answer is correct 23:05 or not because we can match the final answer against the ground 23:08 truth. 23:10 Now, with ORM, we are done there because that's how we 23:14 assign the label for the model. 23:17 But for PRM, the way they did it in this paper 23:22 was that they looked into-- 23:24 they had human annotators go through the steps 23:28 of the generations by the model and assign a score to each step. 23:35 So step 1 in this case was correct. 23:37 Step 2 was correct. 23:38 And probably because, maybe, all of these steps were correct, 23:43 that's why it led to the final, correct answer. 23:46 But basically, they created this entire data set 23:50 of human-annotated, stepwise, correct or incorrect labels. 23:58 And the final reward for this stepwise now, 24:02 again, in this case, could be-- 24:05 what they propose is that can be calculated 24:08 as the product of the stepwise rewards, per step of the model. 24:13 So they basically train the model 24:15 against these stepwise annotations, 24:17 and then they use that. 24:19 They create steps. 24:21 You can even ask the model to generate the answers 24:24 to a solution step by step. 24:26 So the steps are clear. 24:27 They run this PRM against each of those steps. 24:30 They get a score. 24:31 They multiply them with each other. 24:33 They get a range of the quality of the answer. 24:39 And based on that, they can decide whether they 24:43 output the answer or try again. 24:47 So this is called process supervision. 24:50 Obviously, there is this credit assignment, 24:53 and then it's the more precise way 24:55 of collecting data and assigning labels 24:58 to different steps of the problem, 25:00 rather than just looking at the output answer. 25:06 A property here that's very important-- actually, 25:09 this is a really-- 25:11 it could be a pitfall for test-time scaling, 25:14 is that, with process supervision, 25:17 we can manage the false positives 25:19 way better than we can do with just outcome supervision. 25:25 Why? 25:25 Because model might hallucinate. 25:28 And this happens, surprisingly. 25:29 Model can hallucinate and get to a final, correct answer 25:33 while the process for it is really wrong, 25:37 whereas, with process supervision, 25:39 because we are supervising, we are seeing all the steps, 25:43 and we have a score for it, it's less likely that we 25:46 get into the mode where the steps are wrong, 25:51 but the final answer is correct. 25:56 Also, it encourages interpretable reasoning 25:59 and human-endorsed process of solutions. 26:04 It's human-endorsed, again, because humans 26:06 have created the labels. 26:09 So this paper produced this data set, PRM800K. 26:16 It's an open-source data set of 800k step-level labels. 26:25 So the way they did this was a generator model. 26:30 They took a language model, and they 26:32 generate a large number of samples per problem 26:34 in a stepwise manner. 26:37 And then they also did this trick 26:39 of-- because they wanted to collect higher-quality stepwise 26:44 labels, they took into what they call 26:47 convincing wrong kind of answers, 26:53 meaning that they gave higher priority to samples where 26:59 the final answers were correct, but the intermediate steps 27:05 were incorrect. 27:09 And this enabled them to be 2.6 times more data-efficient 27:15 than just randomly selecting these samples 27:18 and give it to humans. 27:20 So they did it iteratively. 27:23 You can have a PRM by just asking an LLM to judge a step, 27:29 right? 27:29 You can have that. 27:31 So they iteratively upgraded the PRM by collecting this data 27:36 and then training the PRM for it and then increasing 27:40 the quality of the data. 27:44 And at each step, the label that was collected 27:47 was either positive or negative or neutral. 27:49 For example, let's take a look at this problem. 27:53 The denominator of a fraction is 7 less than 3 times 27:59 the numerator. 28:01 If the fraction is equivalent to 2 over 5, 28:03 what is the numerator of the fraction? 28:06 And then there are these steps. 28:08 And the first step was labeled correct, second was correct, 28:12 and all the way to the fifth step was correct. 28:14 But this last step, obviously, the model failed to do the math. 28:18 x should be 14, but it's 7. 28:21 So that was labeled incorrect, so that's 28:22 how they collected the labels. 28:29 Here is the process. 28:31 So they took GPT-4 as their base model. 28:35 They still fine-tuned GPT-4 for both the ORM and PRM 28:40 training on this data set first, so the supervised fine-tuning 28:45 first. 28:49 And for ORM, of course, they just 28:52 have the pairs of generated samples and final correctness. 28:56 But for PRM, they had this generator generate 28:59 the sample and then the stepwise correctness label. 29:04 For ORM, the final score was the score 29:07 of the final token in the completion. 29:09 Again, they again went to the token-level approach for ORM, 29:14 but for PRM, it was the product of probabilities for every step. 29:21 Now let's take a look at some of the solutions. 29:24 They were showing that PRM outperforms 29:27 ORM and majority voting. 29:29 So again, here, n is the number of solutions per problem. 29:33 Majority voting, again, fails after, 29:36 in this case, 100 or so samples. 29:39 The blue line is the ORM model, and the orange one 29:42 is the PRM approach that is doing better. 29:49 But they noticed that PRM detects correct solutions 29:52 for some of the very rare occurrences of a correct answer, 29:57 some problems with less than 5% correct answers 30:00 in the distribution of samples. 30:06 They also showed the benefit of larger data in 30:09 and this active learning approach that they had, again. 30:16 And the other observation that they were showing 30:21 is that PRM seems to be more data-efficient. 30:25 So the number of labels that they collected for PRM-- 30:29 if they match that against ORM, both PRM and ORM benefit 30:36 from more and more labels, but it turns out 30:38 that the PRM is more sample-efficient, right? 30:43 So if they have 100 solutions labeled per problem for an ORM 30:48 model, that's like having one solution per problem, based 30:53 on the PRM approach and so on. 31:01 The other interesting property-- 31:03 and that's what we all should strive 31:07 as we are training these kind of verifier and reward model, 31:10 is the generalizability to new domains and new data sets. 31:16 Here, what they're showing is that majority voting actually 31:20 is better than ORM in terms of how it generalizes, 31:24 but PRM is better than majority voting overall, 31:31 and it can tolerate a whole lot more distribution 31:34 shift than majority voting. 31:37 And so, overall, this seems like a much better approach here. 31:43 Any questions? 31:46 Yes? 31:47 So PRMs obviously could lead to false credit assignment 31:51 if a step looks good but doesn't actually 31:53 contribute to the final answer. 31:55 Are there domains that you would expect PRMs to actually hurt? 32:00 So the question is, can PRM hurt sometimes 32:03 because the final answer might still be incorrect? 32:08 So the truth is, a lot of the newer approaches 32:12 combine the two. 32:13 The reality is they combine both the PRM- and ORM-based 32:17 solutions because you want to get 32:19 the benefit of both approaches. 32:22 And the other thing is that the threshold 32:25 that you can define for PRM, then, 32:28 is also a hyperparameter that now you 32:31 have introduced that you have to deal with and optimize 32:34 for your system, and that's an additional complexity here. 32:38 But yes. 32:40 For the previous one, the plot, does it make sense 32:45 to do this comparison? 32:46 Because PRM requests much more labels than ORM. 32:50 So do they show a plot where the x-axis is [INAUDIBLE]? 32:54 Yeah, I think we were trying to. 32:57 Yeah, because they also mentioned 32:58 that it's hard to compare. 33:00 And they did manage to have-- because it's 33:06 true ORM needs k labels for problems, 33:10 but PRM needs k times some number of steps. 33:14 But then it's very hard to control 33:15 that number of steps of-- how do you exactly control? 33:18 But they did try to manage that, and I believe some 33:21 of the graphs-- they do that. 33:23 But that's a good point. 33:25 Yes? 33:25 What does "majority voting" mean in this context? 33:29 So "majority voting" is you have n samples. 33:33 You don't have any access to any reward model. 33:38 You just see which sample or solution is repeated most, 33:42 and you take that as your final answer. 33:44 [INAUDIBLE] the reasoning process, 33:46 so are you also going to match on the reasoning process? 33:49 No, you look into the final answer in this case. 33:54 In this case, are they using the same PRM at 100k for all 33:58 of those tasks? 33:59 Or are they using different terms for each [INAUDIBLE]? 34:03 Because they're talking about generalization, 34:05 I assume this should be-- 34:07 The same? 34:08 The same one. 34:08 But yeah. 34:14 Yes? 34:16 How does this avoid when the reasoning is actually 34:18 bad, but graded high? 34:20 For example, if you go back to slide 23-- 34:25 This one? 34:26 [INAUDIBLE] 34:30 Yeah, this one. 34:31 What if the reasoning is, let's call the numerator x, 34:36 and then x equals to 14, and [INAUDIBLE]. 34:41 How does that avoid [INAUDIBLE]? 34:44 So you were saying the PRM can give a high score? 34:48 What if it's skip all the reasoning? 34:51 Let's call the numerator x, and then x equals to 14. 34:54 How does that avoid that kind of hacking? 34:57 Sorry, I don't get it. 34:59 Can you speak up? 35:00 What's the alternative scenario you are talking? 35:04 The reasoning step would just be numerator x, 35:08 and then x equals to 14. 35:10 And how does the PRM avoid that kind of real [? typing ?] 35:13 scenario? 35:14 I think the model directly skips some steps 35:17 and directly generates the answer, like, [? speeds ?] 35:22 some of the processes. 35:23 So these programs are trained like-- 35:26 programs can also see the previous steps, 35:29 like the question's previous step. 35:30 Here is the current step. 35:32 Score it. 35:33 Right? 35:35 So if x is 14, then it's great, right? 35:37 So it doesn't really encourage the thinking process 35:40 of going to step-by-step. 35:43 It doesn't, no. 35:44 No, it doesn't. 35:45 But you can prompt the model to do this because you're not 35:49 changing the model. 35:50 You're not changing your generator. 35:54 That failure mode could be when you were fine-tune 35:57 your generator model. 35:58 Then the model stops reasoning and just generates 36:02 a final answer that the PRM likes, right? 36:04 But you're not touching your generator. 36:07 The generators can still be prompted to do reasoning, 36:10 create these steps, and then the verifier just scores them. 36:16 Let's say we want to use the PRM to train the-- 36:20 Yeah, then you have to be careful. 36:22 Then you have to make sure that the chain of thought 36:25 makes sense. 36:26 And again, the way you train your PRM 36:31 can also be aware of that because your PRM could be-- 36:37 these are human-generated labels in this case, right? 36:41 If we had skipped all of this, and then we would go to x 14, 36:46 probably, the human says, this doesn't make sense. 36:49 This step is lower, because this step is skipping a few reasoning 36:54 steps. 36:56 Does that make sense? 36:57 So these labels are created by humans here. 37:00 [INAUDIBLE] 37:05 So if the labels are not generated by humans-- 37:08 and we are going to see that in the next paper-- yes, 37:10 that is definitely a caveat and something 37:12 that you can make sure. 37:14 But there are ways to handle that as well. 37:16 So for now, let's assume these are human labels, 37:19 so the processes are labeled in a way 37:21 that they make sense as well. 37:23 We have so many things to encourage these generation 37:27 to make sense. 37:28 First of all, the generator should generate the steps. 37:33 The PRM is trained on human annotations 37:36 that they also make sure that this step makes 37:39 sense, given the prior steps. 37:43 Any other questions? 37:48 OK, let's go through the next paper, Math-Shepherd. 37:53 Verify and reinforce LLMs step-by-step 37:56 without human annotations. 37:57 So we're getting to, what if we didn't 38:01 want to collect that many annotations from humans? 38:06 So again, motivation is that having stronger verifiers 38:14 shows the potential to really improve model's reasoning 38:18 behavior, especially if you do test-time scaling. 38:22 PRMs have shown more potential than ORMs, 38:25 but they require a lot of data, and it's 38:28 hard to collect these data. 38:30 So can we automate the process of label collection for PRMs? 38:37 So this paper talks about two kind of approaches. 38:41 One is automatic annotation, and also brings the reward model 38:50 into the loop of generator optimization via a reinforcement 38:55 learning approach, which-- 38:58 we are going to briefly mention that. 39:00 So, in this case, the way they try 39:06 to go about this automatic annotation of steps 39:11 was that they define the quality of the reasons step 39:16 as the potential of this step to reach 39:19 to the final, correct answer. 39:21 So let's say we are at a given step. 39:24 We sample N completions from that given step 39:28 and see the potential for a correct solution 39:32 at the very end. 39:34 And they introduce two ways of doing this. 39:37 One is the hard estimate, where a step is annotated as a success 39:46 if any of those n generations after that step 39:51 gets to a final, correct answer. 39:55 The soft estimate measures the frequency, 39:58 like what portion of the generations after a current step 40:02 reached the final, correct answer? 40:04 So here is how it works. 40:07 Say here is the problem, and this is our current state. 40:11 Instead of just sampling one step, N is 3 in this case, 40:17 so we sample three steps. 40:18 And then we continue, and we measure the frequency. 40:23 In the soft estimate, we measure the frequency of correct answer 40:27 at the very end. 40:28 In hard estimate, we see whether any of these are correct or not. 40:34 So hard estimate would be 1 for step 1. 40:37 Soft estimate would be 2/3 for step 1. 40:43 Anyone can name a challenge or a drawback of this approach? 40:53 So basically, this soft or hard estimation 40:59 is replacing the human annotations. 41:03 So we're getting a score per step 41:05 by just rolling out more samples from that step 41:08 and seeing whether we reach to the final, correct answer 41:12 or not, and assume that we do have 41:13 final labels for these problems. 41:21 Yes? 41:22 So we-- for unpredictable paths [INAUDIBLE] 41:28 lower when the unpredictable path could be built [INAUDIBLE]. 41:35 Sorry, speak up again. 41:36 If it's unpredictable-- 41:38 Or unusual approaches to solving a problem. 41:42 It could be scored as low or something, 41:44 while that might be the way to be able to solve that problem. 41:48 It might be a really hard problem. 41:50 Now likely to lose that path to that solution. 41:54 That is correct. 41:55 So basically, say N is 3 here. 41:58 Maybe we needed N equal 100 to see 42:02 that final, correct trajectory to score this step, 42:09 but we won't see that if we just look into 3, 42:12 and then we give a score of 0. 42:15 Anything else? 42:23 Yes? 42:24 I guess this is similar, but if you're 42:26 dealing with a hard question, if you're 42:28 taking the approach of generating a lot of samples, 42:31 a lot of the samples can just not be correct or not 42:33 be the correct way to do it. 42:34 So-- 42:36 You don't get any signal for hard problems in this case. 42:42 That's right. 42:44 [INAUDIBLE] 42:48 Exactly. 42:49 I think this is probably-- 42:51 these two that you all mentioned here-- 42:54 for hard problems, you don't get signal. 42:56 The other one is, what if there are wrong steps in the middle, 43:00 and we got to the final, correct solution? 43:03 We still label those steps as correct. 43:08 So the hope is that, if you have more samples, 43:10 it shows the wrongness in one of these trajectories, 43:15 but it's not guaranteed. 43:24 So here's the verification process. 43:27 You sample N again, sample N candidate solutions, 43:32 score them using the PRM, and you select the highest PRM score 43:38 across the generated samples. 43:42 They also took this PRM and trained the model with this PRM 43:47 to encourage the model to generate steps that are scored 43:51 higher and higher by PRM. 43:53 So PRM basically served as the reward model for this. 43:58 So they looked into the hard versus soft annotations, 44:02 and although the loss function as they increased N would be-- 44:09 it seemed like the soft annotations were doing better. 44:14 But somehow what they observed with the N equal 4, 44:21 the depth of the trajectory completions 44:24 to collect these labels-- that's where they get their best 44:28 results anyway. 44:31 So it really didn't matter whether they 44:32 used the hard estimate or soft estimate, 44:35 and they went with the hard estimate at the very end 44:37 because it was just easier to measure that. 44:42 So here is another comparison with baseline. 44:51 "SC" is self-consistency. 44:53 It's another name for majority voting. 44:56 Which solution was most consistent? 44:59 That's the red one. 45:01 ORM is the blue one, and the SHEPHERD is this PRM approach, 45:06 is the green one. 45:07 And the interesting part here was that there 45:10 was no human annotation. 45:12 All the annotations were modeled, generated, 45:17 or the mechanism that they introduced. 45:21 Here, they also compare with PRM800K, 45:24 which was the verifier in the previous paper that was trained. 45:30 And it outperformed that as well on MATH, 45:33 which is like a much harder set of-- this data set 45:37 is more difficult than GSM8K, and they were getting 45:40 an even higher delta in there. 45:46 They also compared these results with different types 45:51 of verifiers again across GSM8K and MATH500 problems. 45:59 And you can see that the difference between these models 46:02 and these approaches-- 46:05 and there is a notable difference 46:10 between these models and the baselines, 46:13 which is self-consistency approaches. 46:18 And this is true across different models-- 46:20 LLaMA-70B, LLema-34B, and DeepSeek-67B. 46:29 They also showed that they can get even better. 46:31 They can make the model even better 46:33 if they RL the model against the PRMs that they trained. 46:37 So, in this case, the Math-Shepherd, 46:39 on their approach-- basically, taking Mistral-7B, this model, 46:46 and then they PPO it against this PRM. 46:50 They're getting much higher results, 46:53 compared to doing this RL with just an ORM. 46:59 The delta here is slightly less, I would say, 47:02 than just the test-time scaling, but this is also another way 47:07 for the models to self-improve. 47:09 You generate the annotation by the model itself. 47:14 You train the PRM and then use that PRM 47:17 to improve the generator model as well, both at test time 47:20 and with this RL fine-tuning. 47:22 So it's multiple levels of using the model itself 47:26 to generate data and RL reward-signal for itself. 47:38 And then they can do RL and PPO and do verification on top, 47:43 and they get even better results, 47:44 but it seems their optimizations are plateauing at some point. 47:51 Any questions from this paper? 47:56 Yes? 47:56 Are there ways to build PRMs that 47:59 incentivize self-error correction, which I imagine 48:02 would be a desirable thing? 48:03 But if we're only rewarding each step being correct, 48:07 we don't necessarily end up reinforcing-- 48:09 you get a lot of incorrect steps back and stuff like that? 48:13 So encourage PRMs to do more self-correction? 48:18 Yeah. 48:19 Yes, absolutely. 48:20 For example, one of the methods that can be done 48:24 is that you have a rubric, or a set of rules, 48:27 that you give the model as your scoring step. 48:32 So the model can check the step against this rubric 48:35 that you have designed for this data set, 48:38 and that can be used as extra signal 48:41 whether this step makes sense or not. 48:45 For example, use a calculator in that case, 48:49 in the equation that we are seeing. 48:51 If the model could do tool use or use 48:54 a calculator in that case, it could solve that problem 48:57 easily-- 48:58 could, for example, use SymPy. 49:02 So there are ways to bring more supervision and more signal 49:07 and enable the model to use tools and test-time scaling 49:11 and agentic approaches to improve the verification 49:14 process. 49:19 Yes? 49:20 [INAUDIBLE] verifiers, do they control the amount 49:23 of test-time compute? 49:29 I believe so. 49:31 You mean "test time" as when they're using the verifier, 49:33 like a baseline verifier, versus themselves? 49:37 I'm actually not sure if they do that when they're 49:40 doing process versus ORM, because, again, 49:42 that's really hard to do. 49:46 [INAUDIBLE] because you're scaling things 49:50 at different rates. 49:52 Yeah, I think that could actually be a question 49:55 that you could investigate. 49:56 You can sample more in the generator. 49:58 You could sample more in the verifier. 50:01 If you're controlling for a fixed compute, 50:04 how do you allocate that? 50:07 So one of the challenges here is that, usually, the ORM and PRM 50:10 should come from the same data for us 50:12 to do a good apples-to-apples comparison, 50:16 because if you do train a PRM, and someone else train an ORM, 50:20 it makes things a bit harder. 50:21 The training data also matters a lot. 50:26 Any other questions? 50:27 Yes? 50:29 Do you could back to the previous slide? 50:31 Yeah, I was just-- oh, no. 50:33 Basically, for this one, I was just curious 50:34 as to why, after using RL, they stuck 50:38 with evaluating performances with greedy decoding. 50:41 I was just curious whether or not raising the temperature 50:45 would have made a difference, or that 50:47 was the only way they could have seen the smallest delta. 50:55 So the question is why they stopped greedy decoding, 50:57 why they didn't do temperature changes and test-time scaling 51:01 again, right? 51:01 I think that's what this slide is. 51:04 So then, in this case, the verification 51:10 is done against 256 outputs. 51:12 So the previous one is no verification, right? 51:16 The verifier is just using the RL training here, 51:18 just at test time as well. 51:25 I think it's something that they didn't do, 51:27 and it would be interesting to see. 51:28 What if they repeat the process again, generate 51:33 another set of labels, and train a verifier, 51:37 but starting from this RL-PPO'd model? 51:47 OK, let's get to the last paper, shrinking 51:51 the generation-verification gap with weak verifiers. 51:55 So this is a paper from Stanford that came out a couple months 52:05 ago at this point. 52:07 So it's a NeurIPS 2025 paper. 52:11 So the motivation, again, is the same. 52:18 We still have a generation-verification gap, 52:21 but the approach here was a bit different. 52:25 Here, we were not training a new verifier. 52:27 We were thinking, can we reduce the generation-verification gap 52:31 by using inference compute and inference scaling, 52:35 and specifically by using an ensemble of verifiers? 52:39 Because when we say "weak," we just assume that no verifier is 52:43 perfect. 52:44 So that's the weakness. 52:45 We didn't purposefully choose bad verifiers. 52:48 These are the best verifiers that are out there. 52:50 But we are using an ensemble of them 52:52 to create a single, much more capable verifier. 53:02 So the property of these verifiers 53:05 is that they do correlate with the true label of the solution, 53:10 but they have imperfections. 53:12 And there are two types of these verifiers. 53:14 We learned about PRMs and ORMs, but another one that we briefly 53:18 discussed a few minutes ago is that we can 53:22 use LLMs themselves as a judge. 53:25 We can show it an answer and ask it, 53:29 do you think this is correct or not? 53:31 The prompt could be literally that. 53:33 Or we can have the model use tools and use rubrics 53:36 and all that to generate a score for a model. 53:39 So that's a whole category of verifiers 53:42 that we can use on top of just PRMs and ORMs. 53:46 So here, this first graph just shows how ensembling and using 53:52 a number of verifiers that are-- they 53:55 could be trained by different labs on different data sets. 53:59 And there are all these language models 54:01 that take a query as input and give it 54:03 a score on the output side. 54:05 Here, we are seeing these four data sets. 54:09 So in this case, we are showing the performance of top 1, 54:14 versus top 5, versus top 10 verifiers, 54:17 based on how good they are in some ranking system. 54:21 Overall, it seems like just ensembling them with each other 54:31 and using them at test time does improve things, 54:36 but not necessarily in a monotonic way. 54:41 But the thing that always helped was that, 54:44 if we train a model, a way for them 54:49 to add them together-- so assign weights 54:52 for different members of the ensemble, 54:54 if we have labels from them. 54:56 Here, we have two methods, Naive Bayes or logistic regression. 55:00 These are very simple methods. 55:02 Basically, you add a single weight per verifier 55:05 to score them together. 55:07 And if you have a training set with labels, 55:13 you can basically come up with these weights, 55:15 and now these results are on the test set, corresponding 55:19 to those tasks. 55:20 So you basically say, for MATH500, I have a set of labels. 55:25 I use that to average my verifiers, find the weights, 55:29 and then I use these averages on test data. 55:33 So it seems like it does give a lot of improvement, 55:38 just this simple task. 55:41 Now, you can do other things. 55:43 So the way Weaver works is by this "score, weight, and select" 55:47 approach. 55:48 So, in the first step, we are going 55:50 to score outputs of the verifiers 55:54 and normalize them so they're all on the same scale, 55:58 and filter out the low-quality verifiers. 56:03 So in this setup, even, we are assuming 56:05 we have access to very limited labels 56:07 that we can say, OK, these verifiers are just 56:09 so bad against our labels. 56:11 We just remove them from our ensemble. 56:14 We have noticed that this is a very important step. 56:17 Your verifiers should be above a certain quality 56:20 to even be let in the pool of verifiers. 56:24 So then you have your verifiers. 56:26 You have your LLM judges. 56:27 You have your reward models. 56:29 And then you get these scores from them, 56:31 whether it's 0 or 1 or the log prob of correctness 56:35 per verifier. 56:38 And then what Weaver did was to use weak supervision to estimate 56:45 the verifier accuracy with very small data, labeled "data." 56:51 And then we use those weights to combine the verifiers' output 56:55 and assign a score. 56:58 I go briefly through how this weak to strong supervision 57:01 works. 57:02 Some of you, if you have taken deep-learning courses, or even 57:05 CS229S last time, we talked about it, 57:11 about weak to strong supervision. 57:13 There's a whole body of work here. 57:15 Weaver was specifically motivated 57:21 by the work by Snorkel, by Alex Ratner et al., and other work 57:26 from Stanford. 57:28 So here is how we have these weak verifiers, like weak signal 57:32 from our verifiers, and combine them together 57:34 to get a stronger signal. 57:36 The setup is the following. 57:38 We have n queries. 57:40 For each query, we create k solutions. 57:43 And then let's assume we have m verifiers. 57:46 So we have a total of n times k times m labels. 57:51 And our goal here is to get this probability of query i, 58:00 sample j equal to 1, given all the labels 58:03 that we get from all the verifiers. 58:05 So how do we estimate this probability correctly? 58:10 The assumption in Weaver is that each verifier 58:12 captures independent aspect of the correctness. 58:17 So this is a key assumption that we made. 58:20 To some extent, the reason we get-- 58:23 so think about it this way. 58:24 If you have a pool of verifiers, if all of them 58:27 agree on the score that they give to any sample, 58:30 then we are not learning anything new from them, 58:33 if all of them are 100% or the same score. 58:36 The signal here is on similarity and dissimilarity 58:41 across this whole set of verifiers with each other. 58:45 So, under this assumption, we can write this probability 58:49 for the response being correct, given the labels that we have, 58:54 and then we use two sets of equations. 58:56 Again, this the first equation is-- 59:00 we can write it this way because we 59:02 assume every two verifier i and j are 59:05 independent from each other. 59:06 The second equation is just the Bayes-- 59:11 the way probability works. 59:13 There's no assumption in writing this. 59:15 And given these two, we can write an optimization 59:22 and come up with the weights that we can assign 59:25 to each of the verifiers. 59:29 And here, we are comparing the results with a naive ensemble 59:34 approach, and it turns out that it does help. 59:37 In some cases, the gains are more significant-- in GPQA 59:41 Diamond and MATH and MLU Pro. 59:45 And each of these problems that are relatively hard problems-- 59:50 for example, the GPQA, our baseline, is already low, 59:55 but the verifier-- this way of weak to strong aggregation 59:59 of responses is helping getting a large boost, 1:00:03 over naive ensembling of them. 1:00:07 Now, there are different ways that we can 1:00:09 scale the verification compute. 1:00:11 We can sample more generations. 1:00:13 So, instead of 10 samples, we can do 100 or 1,000. 1:00:17 We can use larger models for generation and verification. 1:00:21 We can increase the number of verifiers in the pool. 1:00:27 And in general, there are different ways 1:00:30 to assign flops or scale inference 1:00:33 compute for solving a problem. 1:00:37 Here, we are comparing some of these methods with each other. 1:00:41 So the dashed, dark red line is the Pass@K Oracle, 1:00:49 meaning that if we had an Oracle selector, 1:00:54 that would give us the correct answer. 1:00:57 We don't have that, and that's why we 1:00:58 are training these verifiers. 1:01:01 The purple one is Weaver Supervised, 1:01:04 and that's assumed that we have a large body of labeled data 1:01:07 and we can use them to learn these weights. 1:01:10 The blue one is the Unsupervised. 1:01:12 In this case, we're using 1% of the training 1:01:15 labels for each data set. 1:01:18 The orange one is the Naive Ensemble, 1:01:20 meaning we just average them, average the verifiers 1:01:26 without any special treatment. 1:01:28 We still are filtering and keeping 1:01:31 the good verifiers in the loop. 1:01:33 And then all of these are significantly better 1:01:36 than methods such as Majority Voting or Multi-Agent 1:01:40 Verification. 1:01:41 And Multi-Agent Verification, or MAV-- 1:01:47 it's just based on prompting LLMs and asking 1:01:50 them to score a given response, based 1:01:52 on different kind of aspects that is produced, 1:01:56 so a rubric-based, which doesn't do that much 1:01:59 better-- or actually worse, in these two cases-- 1:02:02 than Majority Voting. 1:02:09 But something to notice here is the drastic gains 1:02:13 that the model can have, going from something 1:02:17 like slightly over 40% to over 70% on these hard problems 1:02:25 and, in this case, matching a model like o3-mini. 1:02:29 So the thing that is interesting here 1:02:32 is that we can use this weak to strong supervision to reduce 1:02:39 the gap between model classes. 1:02:41 So what we are seeing here is that, 1:02:43 if we take a generator model of Llama 3.1 8B Instruct, 1:02:48 and the verifier model which is the pool of verifiers that are 1:02:53 8B and below, we get an average of 70% on these data set. 1:02:59 And this is almost comparable to the accuracy that we get with 1:03:05 majority voting, but when our models are at 70B. 1:03:10 So we are making the model, making this 8B model, 1:03:19 roughly the same as the 70B class by just doing this 1:03:24 inference scaling and using the verifier. 1:03:27 And in this case, unlike what we saw 1:03:29 in most of the previous lecture, we 1:03:32 are talking about the end results. 1:03:33 We're not talking about the coverage. 1:03:35 We're talking about the solution accuracy of the entire system. 1:03:41 And here again, in this case, for when we do the same, 1:03:45 apply the same generation and verification to 70B class 1:03:49 of models, we are getting an average accuracy of 86.2% 1:03:53 on these data sets. 1:03:55 And this is very comparable to o3-mini, which has a much-- 1:04:04 which is like a proprietary model and in general, 1:04:07 is a different class of models. 1:04:10 So this shows how much verification, or this work 1:04:13 and verification, can help with the results. 1:04:16 Another work that was done on increasing 1:04:20 the usefulness of something like Weaver 1:04:22 is that a challenge with ensembling verifiers 1:04:25 is that you have a ton of models that now you 1:04:27 need to run for each solution to get a score, 1:04:31 and then to average them. 1:04:34 And that increases the cost, especially 1:04:36 if, instead of one sample per query, you have 100 samples, 1:04:40 and then you want to create all these scores. 1:04:43 So instead, the proposal here is that, what if we do train 1:04:49 Weaver once and then distill it into a much smaller model? 1:04:53 And that's how, instead of running each of these LLM judges 1:04:58 or reward models per sample, we just take this distilled LLM, 1:05:03 which-- in this case, we showed that it could be really, 1:05:06 really tiny, like something like 400 million parameters, 1:05:10 as opposed to the original model, 1:05:12 which was in the 70B range. 1:05:15 And it turns out that, with this distilled model, 1:05:18 we can capture 97% of accuracy of this large pool of verifiers, 1:05:26 but significantly-- in this case, 99%-plus-- 1:05:31 use less compute at test time. 1:05:33 And these distilled version, and the original version-- they're 1:05:38 all open-sourced, and the checkpoints are available, 1:05:42 if any of you are interested in working 1:05:44 with them in your agentic or test-time scaling projects. 1:05:54 And here is another graph. 1:05:55 Here is the distilled one, which is the light blue one, 1:05:58 and the original Weaver is the dark blue one. 1:06:02 And we are comparing, again, the efficiency, the success 1:06:06 rate over the inference compute, the total inference compute 1:06:10 that we are spending on this data set. 1:06:14 And the distilled one, obviously, 1:06:18 is significantly more efficient. 1:06:20 But even the original Weaver-- because it reaches 1:06:23 higher levels of accuracy at those levels, 1:06:27 it becomes more flop-efficient than models like naive ensemble 1:06:30 and majority voting. 1:06:35 And let's do a quick recap. 1:06:39 So today we covered four papers that 1:06:45 also show the trajectory of how this work, research 1:06:49 on verification has progressed over the last four years. 1:06:53 We talked about how verification can improve training, 1:07:00 can improve the outcome and the quality of results, 1:07:04 both during training and inference. 1:07:07 Process reward seems to be very useful and, overall, very 1:07:13 effective, and more so than the outcome rewards, 1:07:17 but perhaps we need to combine them to get the best 1:07:20 results, both PRMs and ORMs. 1:07:23 And we also, in [INAUDIBLE], learned that scaling-- 1:07:27 you need a lot of data. 1:07:28 And the more data you throw in the verifier training, 1:07:31 it becomes better, and you can use it 1:07:33 during your RL fine-tuning to increase the quality 1:07:38 of your generator as well. 1:07:40 And the Weaver work took a whole different approach 1:07:44 to verification, and that was just 1:07:48 doing this weak supervised optimization 1:07:50 on an ensemble of verifiers. 1:07:52 And in a way, we are using test-time scaling, 1:07:55 but by bringing more verifiers, rather 1:07:57 than sampling a single verifier more to get the results. 1:08:02 And it turns out we can distill that and capture 1:08:05 a whole lot of the quality from a much smaller model. 1:08:09 And with that, I can conclude the class. 1:08:13 Any questions? 1:08:15 Yes? 1:08:15 [? Are ?] all the papers were on [INAUDIBLE] reasoning 1:08:21 benchmarks. 1:08:21 Do you think there's anything else that 1:08:23 could work for our project, other ways 1:08:25 to evaluate the models? 1:08:28 Like coding? 1:08:29 Other ways to evaluate the model-- 1:08:31 you mean other areas, like coding, 1:08:34 would be a very interesting area for reasoning. 1:08:38 So there were some coding problems 1:08:41 in some of the benchmarks here, so it's not just math. 1:08:45 But coding is also a very interesting area 1:08:51 for verification. 1:08:52 We will talk about code monkeys. 1:08:55 This is a paper that-- 1:08:57 instead of directly training a verifier, 1:09:00 you can make the model create a test-time system 1:09:03 to make the model generate unit tests, 1:09:06 and those unit tests become your verifier. 1:09:08 So that's an entirely different approach. 1:09:11 But these kind of verifications should also 1:09:14 be useful for any kind of reasoning task. 1:09:22 Yes? 1:09:24 I'm curious, how do generation-verification 1:09:27 [INAUDIBLE] using models that internalize this [INAUDIBLE]? 1:09:37 OK, so the question is, for reasoning model, 1:09:40 how does the generation-verification, 1:09:42 or this kind of test-time scalings work for them? 1:09:45 It actually helps them as well, still. 1:09:48 But the reasoning model-- 1:09:50 we're going to talk about them in this class. 1:09:53 But basically, what they're doing-- a big part of it 1:09:55 is to generate these reasoning steps, use the reward for it, 1:10:00 and do an RL loop, and generate positive trajectories, 1:10:04 and then use that positive trajectories as part of training 1:10:08 for the next round. 1:10:09 So the test-time scaling is inherently used 1:10:13 to generate data and generate the training process. 1:10:17 But they still would benefit because they still 1:10:21 could explore different parts of the solution space 1:10:25 with more sampling. 1:10:27 [INAUDIBLE] existing knowledge, using [INAUDIBLE]. 1:10:34 Do you think, in the future, these two paradigms 1:10:37 will converge, or do you think maybe 1:10:40 repeated sampling will only be used for generating 1:10:44 training data for [INAUDIBLE]? 1:10:48 So the question is, would we converge to a future 1:10:51 where these kind of test-time scaling, repeated 1:10:53 sampling [INAUDIBLE] just for training time 1:10:55 to create the data, and then, at test time, 1:10:59 we just ask the model once? 1:11:01 I think that's the direction we want to go. 1:11:03 We want the model to be really, really good at pass at one. 1:11:07 The first time we ask it to do something, 1:11:09 we want it to generate, and that increases 1:11:11 the efficiency and all that. 1:11:13 That would be the hope. 1:11:14 One challenge, though, with making the model-- 1:11:18 to make the log probs of the model-- 1:11:21 to sharpen through one answer is that the creativity 1:11:25 and diversity of solutions may be lost somehow 1:11:27 during the training. 1:11:30 So surprisingly, we still like models 1:11:37 have a lot of creativity and diversity 1:11:39 in the way that they think. 1:11:41 And that's a good feature that we want to have in the models. 1:11:46 Yes? 1:11:48 Does it matter if generator and verifiers 1:11:50 have very different architectures-- for example, 1:11:52 if the verifier is a mini version of the generator, 1:11:55 versus if they're very different architectures? 1:11:59 Does it matter-- 1:12:00 Which one is empirically better? 1:12:04 OK, so the question is if the generator and verifier 1:12:07 is from the same family or not. 1:12:10 That's a good research question. 1:12:13 I think, in general, it's interesting. 1:12:16 Models like their own generations 1:12:19 and their own way of interpreting 1:12:21 results much better than using a different family of models. 1:12:26 We are going to talk about a paper that 1:12:30 talks about how certain models can benefit from, 1:12:34 like a different class of verifiers over themselves. 1:12:37 But specifically, I am unaware of a study that 1:12:41 has looked into that specific question of same size 1:12:45 and family, versus different size, different family-- how 1:12:50 that changes the behavior of the model.