Title: Part 9 Future Research Areas Authors/Date: Akanksha Bhardwaj, Azalia Mirhoseini — Stanford CS329A Video ID: AyO6wyu4DEg | URL: https://www.youtube.com/watch?v=AyO6wyu4DEg Playlist: https://www.youtube.com/playlist?list=PLangBM27OtEA ------------------------------------------------------------------------ 0:05 So we will start with an overview of what we 0:09 have covered in this quarter. 0:13 So we started the class with just giving you 0:15 an overview of LLMs and how over the past year, 0:21 test-time scaling and train-time scaling 0:23 using self-improvement techniques where 0:25 the fundamental loop is driven by verifiers, 0:27 and feedback from running those verifiers 0:31 and getting rewards, and then using reinforcement learning 0:34 and even search algorithms to help the models to climb 0:38 on math and coding, for example, have 0:41 been a lot of the focus of initial part of the class. 0:45 From there, we started talking about evolutionary strategies. 0:48 So there was a whole class on open-endedness 0:50 where we're not just driving self-improvement loop 0:54 with rewards. 0:56 We also allow the model to just go 0:58 explore if the exploration space can be defined. 1:02 And here we discussed alpha evolve and similar techniques. 1:07 Then we moved to a paradigm where 1:10 we build end-to-end workflows for agents using tool-use. 1:14 And as the models interact with environments and tools, 1:18 you basically are able to drive the workflows end to end. 1:22 And if you imagine that if you have 1:26 tasks that need to be completed over multiple steps, 1:28 then you do need to have perhaps the ability 1:31 to look at knowledge bases. 1:32 So that requires retrieval and memory. 1:35 So basically, as we go towards real-world workflows, 1:39 we need better planning and multi-step reasoning 1:43 in these models. 1:44 So we discussed some works around that 1:47 and how to really look at evaluations 1:50 in the next generation of models and what 1:53 capabilities would drive that. 1:55 And then we had several guest lectures 1:57 on post-training and multimodal agents 2:00 and robotics and reasoning and so on. 2:03 Azalea, is there anything you would like to add here? 2:07 Yeah, I think this sounds good. 2:09 This is full coverage. 2:11 So one of our guest speakers also 2:13 talked about the symbolic techniques, 2:17 which goes under the umbrella of tool use 2:19 and how that could add to the synthetic data, which 2:23 is another trend we are seeing. 2:25 Amazing. 2:25 Thank you. 2:30 So I think just to recap and this-- 2:33 the class is named as self-improving agents. 2:36 So it's not just focusing on LLMs. 2:38 It's focusing on the agent aspect of it. 2:40 What that roughly-- if we were to remind you 2:44 of what we discussed in first class, 2:46 is that how is an agent a generalization of LLM. 2:50 It has a goal and it will go and interact with the environment, 2:53 collect feedback, and use that feedback to correct its steps. 2:56 So they're basically systems that 2:58 can direct their own processes, use tools, and then accomplish 3:02 a goal. 3:03 And oftentimes in the current paradigm, the LLMs by themselves 3:08 are not powerful enough to drive towards a full goal. 3:11 They're getting there. 3:12 So oftentimes we would hand-write these systems 3:16 as workflows where you are orchestrating LLMs and tools. 3:20 However, in certain scenarios like coding agents, 3:24 you are seeing some of these agentic workflows being 3:27 driven by themselves. 3:28 So it's an exciting time where we are seeing 3:31 a lot of progress in this area. 3:34 And how we construct these workflows oftentimes 3:37 means you're orchestrating LLMs, you're 3:39 orchestrating verifiers because you need some form of reward. 3:42 You might have LLMs as judges for verifiers. 3:45 You might have tool calls. 3:46 You might have search algorithms in there. 3:48 So you might have parallel LLM calls that are going and doing 3:52 these kind of search. 3:53 But overall, since an agentic workflow 3:56 is driving towards the goal, It. 3:57 Needs the capability to plan and to reason over multiple steps, 4:02 correct itself if it were going in the wrong direction, 4:05 and overall keep improving in its capabilities, which 4:10 is where the self-improvement aspect comes in. 4:13 So in this lecture, as we wanted-- 4:15 we wanted to cover some future research areas. 4:18 So we organized it as key ideas from certain papers 4:22 that both cover the self-improvement side of things 4:25 and just generally the efficiency side of things 4:29 for intelligence, both of which we believe 4:31 are extremely important directions in the next research 4:36 areas. 4:38 So for the first three papers that I will cover, 4:44 I'll just present the key ideas. 4:46 But let me just present why those papers were chosen 4:50 and what are the key ideas that they point in terms 4:55 of next steps and directions. 4:57 So what we learned in the class is a lot about say, 5:00 test-time scaling, where you could improve 5:04 by having multiple samples and majority voting and so on, 5:08 and then train-time scaling where 5:10 you can use that feedback in the inference loop 5:13 and get rewards to drive that self-improvement loop. 5:17 Oftentimes, this is still limited to narrow domains 5:21 like math and coding. 5:22 So how do you get generalization across these domains, 5:25 and how do you generally get those reasoning chains 5:28 to be diverse enough so that you can continue 5:31 to drive that self-improvement loop is something 5:34 that is an important research area even now. 5:38 An important aspect of building these self-improvement loops 5:43 continues to be verification. 5:44 So how do you have robust verification techniques or meta 5:49 verification techniques where you 5:50 verify what was going into the reasoning 5:53 chain is often quite valuable. 5:56 So that's another area that the second paper I will present 6:00 the ideas on will cover. 6:01 And the third aspect is that even though we 6:04 talk about train-time scaling, the prompts that 6:07 go into training these models gets selected very statically 6:12 and require humans to select them. 6:14 So how do you break through the data barriers 6:17 so that the self-improvement loop 6:19 can pick the right set of data that drives that loop? 6:24 So that's going to be the first part of the focus. 6:26 And the second part of the focus will 6:28 be focused on just what drives efficiency for intelligence, so 6:33 intelligence per what aspects. 6:35 So let's take a quick look at some 6:37 of the ideas that might drive this self-improvement loop 6:41 further in the next generation of systems. 6:44 And given how fast the field is moving, maybe 6:47 we'll be teaching papers on each of this next time. 6:51 So the first paper that we cover is 6:54 coming from multi-agent finetuning, 6:55 which focuses on self-improvement 6:57 with diverse reasoning chains. 7:00 The key idea of this paper is that pre-training, compute, 7:05 and instruction tuning is bottle-necked by-- 7:07 so pre-training compute is a compression of internet scale 7:10 data and instruction finetuning often requires real human data, 7:13 where you need human preferences to decide what is good model 7:18 response and bad human response-- 7:21 bad model response. 7:22 So an alternative to this is that you 7:24 can use the data generated by LLMs, which is often 7:27 called as synthetic data. 7:28 And you can do iterative finetuning 7:29 where you generate possible solutions 7:32 and filter out the incorrect ones. 7:33 This is called rejection sampling and fine-tuned 7:35 only on the good ones. 7:37 Now there are many variants, STaR 7:39 and so on, which can help drive this loop even better. 7:43 We covered STaR in class. 7:45 You can even add reasoning chains in that process. 7:49 But what typically happens is that if you're 7:51 using a single large language model to generate the data, 7:55 then it will generate solutions that will be very similar. 7:59 And oftentimes the performance increase 8:02 will stop after a few iterations or after tens of iterations. 8:07 And typically, this problem comes down 8:11 to really the notion of lack of diversity. 8:13 So when you look at model training, 8:16 at the pre-training scale, the data 8:17 is so diverse because it was generated over such a long time 8:21 by humans. 8:22 So it is quite diverse and that helps the performance. 8:27 But when a single large language model is generating outputs 8:31 for a set of prompts, it will not have very diverse responses 8:34 even at high temperatures. 8:36 So the intuitive solution for this 8:38 often means that you can use multiple agents, which 8:41 are specialized in some way, and then use 8:44 them to improve the diversity. 8:47 So what this paper proposes, for example, 8:50 is multiple specialized agents for generations 8:53 that will then produce a diverse initial solutions. 8:56 So they train generation agents and critic agents. 8:59 So the generation agents will help generate 9:03 a lot of diverse initial solutions, 9:05 while the critic agents will evaluate and refine 9:07 the solutions. 9:09 So those are specialists as well. 9:11 So if you were to look at step-by-step process, 9:14 and we don't need to go through every single detail 9:17 in the algorithm below, but the generation agent 9:20 will come up with an initial answer 9:23 and then there will be debate over multiple rounds. 9:26 So what the generation agent will-- 9:29 in the initial step, it will just have the initial answer, 9:32 but in the subsequent steps, it will summarize and get 9:36 a summary of all other agents responses 9:38 and use that also as part of generating the next response. 9:43 The goal of the critic agent is to critique 9:46 this updated set of answers. 9:47 So basically, if you're doing any iterations, 9:50 there's an initial response from each of the generation agents. 9:55 Then you do a summary across all of these agents, 9:57 and then the critic agent will critique 9:59 this updated set of answers. 10:01 So it's very similar to what you guys were already familiar with. 10:05 The main difference here is that instead 10:07 of getting the critique over a single agent's response, 10:13 you're actually using multiple agents responses 10:15 and then specializing and summarizing across them. 10:19 What this enables is that you have some form of diversity, 10:24 even at the generation stage, before you're 10:26 critiquing the answer. 10:27 So you basically get majority voting for free 10:29 just by having multiple agents and assuming that these are 10:34 trained slightly differently. 10:36 You get diversity for free. 10:39 So the generation models in this particular case 10:42 are fine-tuned from the same base model. 10:45 And the goal there is to produce good answers given a question. 10:49 So over multiple iterations, what they're doing is 10:52 they're taking outputs and filtering 10:54 for a match with the majority vote. 10:56 And then doing SFT on this list of prompt and response pairs. 11:00 If you were looking at a very naive version of doing this, 11:03 you could just use different models, 11:05 which is something that people do in practice as well. 11:08 So they will get very different responses 11:10 with different prompts. 11:11 So that's the poor man's version of doing multiple generations. 11:17 And the critic models in this particular paper 11:20 was proposed as the critic models will basically 11:24 take the updated answer and then select the best one. 11:30 So they are also fine-tuned. 11:32 And these are fine-tuned on a mix of trajectories where answer 11:36 is correct at the start and then is corrected 11:38 over a course of debate. 11:40 So basically, the critic model is 11:42 learning how to contrast the correct and the incorrect 11:47 answer. 11:49 So to summarize the whole process, 11:52 there is n generation models and they are each 11:55 producing an output. 11:56 The summarization happens. 11:58 You can imagine that either you summarize it with a model, 12:01 or you can just concatenate the responses of all the models. 12:04 Then there is a critique process on that 12:08 and that's added to the input. 12:12 And then you do a second round for all of the model 12:15 generates outputs, updated answers, and then 12:18 you do majority voting on top of that 12:20 and again summarize the responses. 12:23 And you can continue this process 12:25 in form of a debate in some ways. 12:27 So those trajectories can then be 12:29 used to fine-tune different critiques as well. 12:35 So how does this help? 12:37 Why did we discuss this? 12:38 So I think the key metric that we're 12:40 looking for is that on the x-axis, 12:44 you're looking at how many iterations of fine-tuning 12:47 it has gone through, and on the y-axis, 12:49 you either have the negative log likelihood 12:51 or you have the embeddings dissimilarity. 12:54 The negative log likelihood is just a proxy for performance. 12:58 And the embedding dissimilarity, the higher the value, 13:04 the more diverse it is. 13:07 So what we see on the left is for two open source models. 13:11 When they go and do fine-tuning, they're 13:13 actually able to continue to increase performance 13:15 when they do this multi-agent fine-tuning. 13:17 And Llama seems more responsive to this technique. 13:22 I think the right side is more interesting. 13:23 As they're going through this process, 13:27 the responses continue to stay quite diverse. 13:30 And this is over math as a data set. 13:33 And overall, they tried this again 13:36 in math over three open source models. 13:39 And since they were going with fine-tuning, 13:41 they had to stick with open-source models. 13:43 But with multi-agent techniques, they 13:45 were able to continue to see improvements even 13:48 over multiple steps of fine tuning. 13:50 While for single-agent fine-tuning, 13:53 the accuracy collapses or doesn't continue to improve. 13:58 And they also show that not just in domain, which is math, 14:03 in adjacent domains like GSM 8-k, 14:06 this is a slightly older piece of work, 14:08 so the numbers are not that high. 14:10 But even in adjacent domains like GSM 8-k, 14:14 the fine-tuned agents actually show higher performance. 14:18 So this technique actually does help, 14:22 so it generalizes beyond just in domain data 14:25 sets that it's trained on. 14:27 So that's the first set of comments. 14:30 So what that roughly says is that if you want 14:33 self-improvement, the reasoning chains that 14:35 are provided to the model to drive those 14:37 need to be diverse in some way. 14:39 And how you generate that one technique could 14:41 be based on having multiple models or multiple agents 14:44 generate those kind of chains. 14:47 A second problem that's very challenging, 14:50 and we covered actually in the first few lectures 14:53 is around verification. 14:54 So it's quite hard to verify model outputs. 14:58 And typically, you would just look at the outcome. 15:03 So if you have a math problem or a theorem proof, 15:08 what you will look at is the final outcome. 15:10 And then you have an outcome reward model, 15:12 so does that match the ground truth, 15:14 and then you will say that this verifies. 15:16 One of the questions that we discussed in the verification 15:20 lecture is that what happens if the reasoning chains are 15:23 wrong, and will that lead the model astray, 15:26 or will that cause some performance gap at the end? 15:31 And we also realize that process reward models 15:34 are quite hard to build. 15:35 So DeepSeekMath-V2 is a recent paper 15:38 that came out which tries to show 15:40 that, at least in the theorem proving domain, 15:42 it is possible to build self-verification loops, which 15:46 are automated. 15:47 So that's a very interesting concept. 15:51 Just to remind you on what I was saying, 15:54 the current reinforcement learning approach 15:55 assumes that final answer will match the ground truth, 15:59 and that's how you're basically creating the rewards. 16:02 So this has enabled saturation of multiple benchmarks 16:05 like Amy or other math-related benchmarks. 16:08 But oftentimes, even when you have the correct answer, 16:11 you might not have the correct reasoning. 16:13 And if you are actually trying to get the models to do better 16:16 in theorem proving, you do require 16:18 rigorous step-by-step derivation, 16:19 which the final output doesn't quite give you. 16:23 Now, the interesting aspect is that the large language models 16:26 are often trained on quantitative reasoning, 16:29 so the proofs that they might generate 16:31 are mathematically invalid. 16:33 And if you ask them to verify, they 16:37 will claim that the incorrect proofs are valid. 16:39 So LLM as a judge technique doesn't quite work here. 16:43 But if you actually have expert humans look at these proofs, 16:47 they will be able to look at the proof and reason about the fact 16:52 that they know this area. 16:53 They can reason about the fact and say that, OK, 16:55 this proof has issues. 16:56 It's just not making sense because this next step is not 17:00 following from the last step or there's reasoning gaps 17:05 in the proof solution. 17:07 So what DeepSeekMath-V2 proposes is 17:11 that instead of just training the reward models, 17:15 they train verifiers or meta verifiers basically. 17:19 So they get humans to identify issues in the proofs 17:23 without any reference solutions. 17:25 And based on that, they train LLMs 17:28 to identify issues in the proof and improve upon those. 17:34 And then if you will find that there-- so 17:37 you basically have trained a model which can now 17:39 critique the proofs themselves. 17:41 So what the final architecture for DeepSeekMath-V2 looks like, 17:46 you folks are already familiar with the fact 17:47 that there's a generator and a verifier, 17:49 typically in the loop for building train-time scaling 17:53 system. 17:54 So you have one model that is producing proofs 17:57 and another model that is perhaps verifying things. 18:01 So it's identifying issues. 18:03 What DeepSeekMath-V2 adds is this notion of meta verifier. 18:06 So you have a meta verifier that will 18:10 review the verifiers analysis for whether it makes sense 18:13 or whether there are issues with the proofs. 18:15 And then the verifier is trained to take these issues 18:19 and then score the proofs on a scale of 0.5 and 1. 18:23 So the verifier will then go improve 18:25 the generator and the generator will 18:27 produce harder proofs, which will then 18:29 go improve the verifier. 18:30 So basically this ends up being a loop in some way. 18:35 And the other interesting aspect here 18:37 is that because once you see the data 18:40 for finding incorrect proofs, then you can actually 18:44 automate that kind of labeling. 18:46 So you don't have to just rely on humans identifying 18:48 issues in the proofs once the meta verifier starts 18:52 to learn this. 18:55 So this notion of meta-verification 18:57 is generally quite interesting in that 19:00 verifiers can get correct score when the reasoning chains are 19:04 incorrect. 19:04 For example, they might come up with fabricated errors. 19:09 So meta-verification has evaluation of this analysis 19:13 to these issues that the verifier-- verifier here 19:17 is LLM as a judge, or do these identified issues actually 19:21 exist? 19:21 Does the score follow from the issue? 19:24 So they're basically analyzing whether the verifier 19:26 did a good job or not. 19:27 And then they have experts annotate 19:29 the quality of this evaluation. 19:32 And just by adding this additional block, 19:34 they are able to improve the quality. 19:36 So it's almost like you had reasoning chains 19:39 and then now you have the verifier over those reasoning 19:43 chains, which is basically saying that, 19:45 is this evaluation correct or not? 19:48 So there's an LLM as a judge, which is identifying the issues 19:53 in this reasoning chain. 19:54 And then do those issues exist or not, 19:57 is what meta-verification is looking for. 20:02 If you look at the results here, just 20:06 with this iterative optimization, I think all of us 20:10 are familiar that a lot of the RL loop 20:12 often is built on top of TRPO. 20:15 In this particular case, they're building on top 20:17 of DeepSeek-V3-based models. 20:19 So just the verification generation loop now 20:23 benchmarked on IMO problems and CNML problems, 20:29 the Gemini version of that was presented 20:31 by one of our guest lecturers. 20:34 But this is an open-source version 20:35 that was released recently. 20:38 They are able to show that in eight iterations just past one, 20:43 the proof score continues to climb just 20:45 with that iterative loop. 20:47 And then if you have best at 32-- best at 32 20:49 means you pick the best solution out of 32 generated proofs, 20:54 there you are almost getting to 42% in proof score 20:58 for IMO shortlist of 2024. 21:01 So this is quite promising as a hill-climbing technique, which 21:05 roughly suggests that there is some promise if you can identify 21:09 the issues in the reasoning chains 21:10 and actually push the model or nudge the model 21:12 towards correcting them. 21:15 So overall, the insights here are that the best proofs will 21:21 achieve higher verification scores, 21:23 and the generator can learn to differentiate higher quality 21:27 proofs from flawed proofs. 21:29 And the self-verification as a result 21:31 can basically have this better improvement loop. 21:36 So verification can be a bottleneck 21:38 and this is one way to break that bottleneck. 21:41 Overall, I think if you were to try 21:44 to generalize this technique in other areas, what you're 21:47 basically looking for is if you have LLM-based verifiers, 21:51 you need to find a way to get them to identify issues 21:54 without reference solutions. 21:56 That's what this paper is relying on. 21:58 Then you need some form of an additional block, which 22:00 is meta-verification so that it can reduce 22:04 the chance of hallucinated issues 22:06 that were identified by the verifier. 22:08 And then you add an additional incentive 22:11 for the generator to maximize quality 22:15 through deliberate reasoning. 22:17 So you're basically improving the quality 22:18 of the model responses here. 22:20 So overall, this notion of self-verification 22:23 is introduced by our DeepSeekMath-V2. 22:26 And it starts to break the bottleneck in verification 22:31 if you can make that more automated end to end. 22:35 Of course, this is still limited by domains 22:39 where verification is easier rather than more difficult. 22:47 So we covered through what is the diversity in reasoning 22:52 chains and we discussed a bit about the verification 22:58 bottleneck in the models where oftentimes 23:01 to get the models to get over the bottleneck of verification, 23:07 you would need some form of verification techniques. 23:09 And DeepSeekMath-V2 publishes that in open source. 23:14 A third problem that you have to solve 23:15 is this notion of what data or what prompts should I train on. 23:23 So there's this very interesting paper 23:25 that has not seen much use yet, but is 23:30 quite promising and interesting is that you basically 23:35 have the models propose the set of tasks 23:38 that they should go train on that are just 23:40 at the edge of what they have learned. 23:43 So to remind you, typically in the current AI training stack, 23:48 you are either using human-curated reasoning traces, 23:52 so that's for the supervised learning stage, 23:55 or for reinforcement learning with verifiable rewards, 23:58 you are expecting some experts to curate 24:00 the question-answer pairs, but that basically requires 24:04 that if you're constructing such a model in math, 24:08 then you need math experts. 24:09 If it's an IMO problems, then you need IMO experts, 24:12 or if it is encoding, then you need strong software coders. 24:17 So as the models continue to surpass human intelligence, 24:22 the ability to find more and more experts and more and more 24:26 such tasks starts to be limiting. 24:30 So what this paper proposes, and this idea is still very new, 24:35 is that a single model can both propose tasks and then solve 24:39 them. 24:39 So it almost goes to the other extreme of, 24:42 we should not really need an external source of data. 24:48 So basically, we should not need human-generated prompts 24:54 to climb on. 24:55 The model itself will propose the tasks 24:57 and then it will go solve them. 24:59 And so that's a very interesting set of ideas. 25:03 And it is feasible in certain domains. 25:06 So they focus more on coding as a domain. 25:08 And they effectively construct three types of tasks-- 25:13 abduction, deduction, and induction. 25:16 And I will explain this diagram in the subsequent slides. 25:19 But effectively, what this is saying 25:21 is that in the proposal stage, they 25:24 will construct the tasks based on this task types 25:28 and then come up with a reward to select 25:30 which tasks make sense. 25:32 And in the solution stage, they will verify and then have 25:36 an accuracy reward. 25:37 And then that can help give a joint update on the model 25:40 itself, as opposed to just one-sided updates. 25:45 So what is the proposer? 25:48 How is the proposer proposing tasks? 25:49 That's quite interesting. 25:51 So it's effectively based on coding paradigms. 25:53 So it's effectively saying that you have deduction tasks where 25:58 you are generating a program and input 26:02 and the environment is executing to get output. 26:06 So this is your usual come up with a program and an input 26:10 pair. 26:11 And then the environment will execute to get outputs. 26:15 Abduction tasks-- they're similar to deduction. 26:18 You're generating the program and inputs and environment again 26:21 computes outputs, but induction tasks are different. 26:24 It samples an existing program and then generates 26:27 new inputs for it plus a natural language message describing 26:31 the function, and the environment then 26:34 goes and executes and decides whether this is going 26:39 in the right direction or not. 26:43 And another interesting bit that I should mention here 26:45 is that the proposer will be conditioned on past here. 26:49 So I'll cover that in a later slide 26:52 because otherwise it will get confusing. 26:54 But it will be basically conditioned 26:56 on past-generated examples that are explicitly 26:58 prompted so that they're basically 27:02 added to promote diversity. 27:05 So how do you select tasks? 27:08 The proposer will get a reward based on optimizing 27:12 for task difficulty. 27:13 So if when the solver can-- so each of the tasks 27:18 that are generated are passed to the solver. 27:20 And if the success rate is zero, then you get a zero reward. 27:26 And if the success rate is non-zero, 27:28 then you basically get 1 minus average success rate. 27:31 And the kind of task that you want to select 27:34 are the ones that are not trivial 27:36 and the ones that are not impossible. 27:38 So there is some ability to continue to learn. 27:40 So you end up selecting tasks that are of moderate difficulty, 27:44 where solver will sometimes succeed and sometimes fail. 27:47 And this will generally help with getting the model to learn. 27:53 So that's the goal of the proposer, 27:56 is to generate tasks for optimal task difficulty 28:00 at a current set of model weights. 28:02 And as the model becomes more capable, 28:04 then the proposer should learn to propose harder problems. 28:09 Of course, if you generate tasks, 28:11 then you need to somehow make sure that these are valid tasks. 28:14 So in the coding abstraction, one can run program integrity. 28:18 So you can actually execute these tasks 28:20 and see whether there are any errors. 28:22 You can do safety checks. 28:23 You can also make sure that if you run these inputs 28:27 multiple times, you get exactly same outputs because 28:29 in code that's possible. 28:31 So the proposed tasks before it goes into the training pipeline 28:35 is validated. 28:36 So that's one way to make sure that the proposal doesn't just 28:42 hack and come up with garbage tasks. 28:45 And in terms of-- there's a buffer that is kept. 28:48 So for every seed triplet, there's 28:52 an identity function like basically the input-output 28:56 and the program for each of those triplet tasks, 28:59 you're adding them to the task buffer. 29:02 And the proposal can sample references from that buffer. 29:06 So there is some sort of a buffer that is kept along 29:09 with how many times the model is succeeding or failing on this. 29:16 So this is basically leading to this idea of curriculum 29:19 learning that is evolving over time. 29:22 And why is this exciting or interesting? 29:25 I mean, on coding benchmarks, this notion starts to let you 29:29 hill-climb despite having much human curated data which-- 29:35 humans can generate a lot of coding data, 29:36 but that after a while, synthetic data or some-- 29:40 this is almost a play on synthetic data in some ways, 29:43 but it's generating that in the loop at the task level as well. 29:48 So what they show is that they're able to get 29:51 state-of-the-art on coding benchmarks, 29:52 even though they didn't have any human-curated data on the prompt 29:56 side. 29:56 And they are able to outperform models 29:59 trained on tens of thousands of expert examples. 30:02 Some emergent behaviors that they 30:04 show are that the complexity metrics increase over time 30:07 so that can be expected if the proposer is continuing 30:10 to increase the difficulty of what it's learning. 30:13 And another aspect is that the diversity 30:16 of programs and answers, when the loop is set up correctly, 30:18 is improving. 30:19 And the proposer is actually increasingly 30:23 is generating more and more difficult tasks 30:26 as the training is progressing. 30:28 So in some ways, it's almost like game theory 30:31 where the proposer and solver are slightly adversarial, 30:34 but overall, they are helping each other improve in some ways. 30:38 So this is solving to the bottleneck of which tasks 30:41 to go train the models on. 30:43 So if you were to look at takeaways here, 30:46 I think it's basically going down 30:49 the path of the task selection problem 30:51 on which these models run their self-improvement loop, 30:55 should not get bottlenecked on human data, 30:57 because otherwise you are very bottlenecked on which 31:00 human experts, as long as you can 31:02 have some form of verification there. 31:04 And you need some form of ability 31:06 to do curriculum learning. 31:08 So this particular technique definitely relies on that. 31:12 And even though they only hill-climb on self-proposed code 31:17 tasks, they actually see strong performance 31:20 on both coding and math benchmarks. 31:23 So that's actually surprising that they are also 31:25 seeing strong performance on math benchmarks, 31:27 though one can explain that from other results that 31:29 have been seen in the field. 31:30 And overall, larger models get bigger gains 31:33 relative to smaller models. 31:34 So this is another idea in the space 31:38 of improving the self-improvement loop that 31:41 seems quite promising and continuing 31:43 to hill-climb in that area. 31:46 Can I add something on the previous? 31:49 So I feel like we are seeing the same thing. 31:52 Like, we learned about SWiRL, the step-by-step RL 31:55 with synthetic data. 31:57 I feel like the findings here and there really 31:59 kind of resonate with each other in the sense 32:02 that synthetic data like model generating its own training data 32:08 can improve the model not only on the task 32:13 that the generation is happening on but also in transferability. 32:17 Like in SWiRL, it was on, let's do a multi-hop question 32:21 answering with retrieval. 32:22 But then as more synthetic data was generated 32:25 and RL training was done, the model 32:27 was getting better at other tool calling, like calling a Python 32:32 and solving math problems, and vice versa. 32:35 So it seems like there is this clear kind of repeated trend 32:39 that we are seeing in terms of synthetic data generation 32:42 by the model and generalization. 32:44 And the other piece that's interesting, again, 32:46 that kind of is the same kind of observation 32:50 is that larger models seems to be better 32:53 at absorbing this kind of data flywheel 32:56 and also generalizing through that, especially when it comes 33:00 to in SWiRL at least on the RL optimization side, which 33:03 is an interesting observation. 33:07 Great, great. 33:09 Where would you take it from what you learned as well. 33:13 So I was going to summarize and you can add more bullet points. 33:16 I think to me, to achieve true generalization and reasoning, 33:20 we really need that notion of diversity in the reasoning 33:24 chain. 33:25 So that was the first set of comments I was making. 33:28 And how we achieved that continues to be an open problem. 33:32 We need better verification, and the more 33:34 we can break that verify loop in a way 33:37 where we are not bottlenecked by humans or tasks that 33:41 are verified only through human experts, 33:45 the better off we are in some ways, 33:47 like building reward models is challenging. 33:49 And then I think the third is the selection of data itself. 33:53 There's only so much data that can be curated by humans 33:57 for what prompts go in. 33:58 So breaking that barrier is also quite 34:00 important in next generation of self-improving models. 34:05 What else you would add there? 34:08 Yeah, I think, one, just like based on these kind of existing 34:13 work, something that would be interesting 34:18 is that how much we can push the frontier 34:21 of this self-improvement only based on the verifiable domains 34:26 that we can generate synthetic data 34:28 and bring it back to the model on diverse tasks, 34:32 on diverse, verifiable tasks, how much that pushes 34:35 the performance and generalizes to other domains 34:39 when we don't have automated verification. 34:43 Like, can we generally make the model smarter and smarter 34:46 in areas that we can verify and minimize or remove 34:55 the need for labeled data in on non-verifiable domains based 35:01 on that? 35:01 So I think that would be an interesting research 35:04 area, of course, a very compute-heavy research 35:08 experiment. 35:11 There are some questions that people 35:13 are asking-- what are some other applications of AI that don't 35:16 fall into verifiable problems? 35:21 So it comes in-- 35:22 I can just think about some areas 35:27 that verification is very slow, like scientific discovery, 35:33 running a very slow simulation in chip design that 35:36 takes like a few days to run to collect one kind of reward 35:41 signal, or running this chemical experiment that you actually 35:45 need to go to a vet lab and run things and see 35:49 what the output, the quality or the reward is. 35:53 Those are the domains that the verification is really hard 35:57 because in RL fine-tuning, or in test-time scaling, 36:03 we need these verifiers to be almost instant, 36:05 or we can wait a little bit. 36:07 Maybe we can wait minutes. 36:08 Maybe we can allow one hour. 36:11 But if in this RL training route-- 36:13 RL training, we need like hundreds or thousands of steps 36:17 of iteration, we can't wait like days or someone human 36:23 in the loop to collect the reward for us. 36:27 So those are the non-verifiable domains-- 36:31 those are some examples of non non-verifiable domains. 36:35 I think truly non-verifiable also comes 36:37 from creativity or any metrics where there's subjectivity 36:41 in some ways. 36:44 Like creative writing, it's like designing an exact reward 36:50 function for that is hard. 36:52 But of course, we can always model that. 36:54 But then you can-- the RL or the agent 36:58 can do reward hacking if the model is slightly off base. 37:04 Any other questions from the class on this aspects? 37:11 Yeah, we have one question. 37:12 Can you hear us? 37:13 Yes, we can hear you. 37:17 You mentioned the harder areas to verify, 37:21 like how are people attempting to sort that? 37:28 For example, like chip design or bio disturbance? 37:37 So one way to go about it is to train 37:43 a model, like a separate reward model that can predict 37:47 the outcome of a simulation. 37:49 So instead of actually running those expensive simulations 37:54 or emulations or through experiment, 37:56 you collect a lot of data on them offline, 38:00 and then you train a reward model that for a given input, 38:04 it can predict the quality of the output 38:07 and you use this reward model as your verification object, 38:12 like in the loop of RL optimization and so on. 38:16 But of course, the generality of the reward model 38:19 is a function of how much data you have. 38:21 And it can be inaccurate and that can cause problems. 38:28 So that's one area. 38:30 Akanksha, do you have any? 38:32 So I think we talked about chip design and so on. 38:35 But even before going to chip design, just the kernel bench 38:38 stuff that was done at Stanford in your lab, 38:41 even there, for example you get compiler execution. 38:46 So you can see that, did you generate the right code or not? 38:49 You can get some metrics. 38:50 But there's a class project that is actually 38:53 trying to read performance profiles as you 38:55 make things more complex. 38:56 And that's a harder problem. 38:59 And so you basically, even though you 39:01 might be able to do design-optimized kernels 39:03 for simple enough things, but if you actually 39:05 concatenate all of these things together 39:07 and you want to performance profile and change 39:11 parts of the code that are not optimized, 39:14 that's a harder problem in some ways. 39:17 And oftentimes the way people will go about solving it is OK, 39:20 the models can solve maybe the kernel optimization problem. 39:25 So maybe I break down the problem into subparts 39:28 and then get the model to look at each part 39:31 and then have some reference solutions to look at a knowledge 39:34 base. 39:35 I'm literally describing A class project right now. 39:37 So what I'm trying to say here is that because of this 39:42 being a hard-enough problem, and the models 39:44 not being at that capability, often it 39:46 requires breaking down the problem in interesting ways 39:49 to what the models are able to do now. 39:59 Let's move to a more along the lines of, OK, we 40:03 can do the self-improvement. 40:04 But you're now doing a lot more inference. 40:07 And there's a cost to this intelligence. 40:10 So Azalea will cover more on how to get efficiency there. 40:14 Sounds good. 40:15 Should I present myself or would you scroll? 40:19 Like what would-- 40:21 Whatever is fastest, I can continue to. 40:23 If you just tell me next, I'll continue. 40:25 OK, sounds good. 40:26 So next, please. 40:28 So this is a recent work that we did in collaboration 40:33 with Professor Ray and Professor John Hennessy 40:37 and a large team of great collaborators. 40:42 So in this project, what we are looking 40:45 for is like looking into the trends 40:48 that are happening on the model side 40:51 and also on the hardware accelerator side. 40:54 And what we are interested in thinking about the future 41:01 and thinking about how the workloads and AI inference is 41:07 going to look like in terms of the type of the size 41:11 of the models and the type of accelerators 41:13 that we are going to need to serve them. 41:16 Next please. 41:20 So right now, we are in the mainframe era, 41:26 meaning that the LLMs that we are running 41:32 are using ChatGPT, Gemini. 41:35 A lot of these larger models that we are running, all of them 41:38 are being run on cloud. 41:42 We don't run them locally because these models are large. 41:46 Some of them we even don't have access to them. 41:48 They're proprietary and even the open source fund for the larger 41:52 models, we are still using cloud to run these models. 41:55 And at the same time, what the observation 41:58 is that the demand for compute as a result is exploding. 42:03 Like Google Cloud grew in something in the order of 1200x 42:08 in the last 20 months in terms of the compute serving. 42:13 NVIDIA had a 10x year-over-year kind of growth. 42:18 And this is crazy. 42:20 These numbers are explicitly driven by AI. 42:25 And as a result, we are going to need something 42:28 in the order of 250 gigawatt of data centers 42:33 to be able to serve this demand right now. 42:35 And again, this is exploding. 42:37 So that means you're going to need more and more energy 42:42 kind of supply for these data centers. 42:45 On the right side, we're showing this graph 42:48 of the number of tokens that are kind of from on the Google 42:54 side has been processed. 42:56 So in the February of last year, it was 160 trillion. 43:00 And in October of this year, it was 1.3 billion. 43:06 So this is one of the fastest growing compute demands 43:10 that we are seeing in history. 43:12 Next, please. 43:15 At the same time-- so we looked into what kind of workloads 43:19 are, or the type of activities or AI 43:22 serving demand that we are observing. 43:26 And then from this data set, this very large-scale data set 43:31 of ChatGPT users, it turns out that something in the order 43:35 of 77% of requests are for task like practical guidance, 43:42 information-- 43:44 like asking for information or writing. 43:47 And a lot of these tasks, it turns out 43:50 that we don't need the very best frontier models to answer them 43:55 correctly, rather smaller and local models 44:00 can be used to address them. 44:03 So on the right side, you're seeing this kind 44:05 of categories of user queries, the type of user queries 44:10 that exist over time. 44:13 And what is interesting here is that, of course, 44:17 users are going to ask more and more complex 44:19 problems from the chatbots over time because the chatbots are 44:23 getting better. 44:24 But still, it's kind of like the vast majority 44:28 of the type of queries that are being asked 44:31 are on the side of, again, the simpler side of complexity that 44:37 can be addressed by smaller models rather 44:41 than large proprietary models. 44:44 Next, please. 44:48 Another trend that is happening right now 44:51 is the improvements in local inference accelerators. 44:56 So the graph on the right, what it shows is that since 2012 up 45:01 until now, we saw something like 126x improvement in the GPU 45:08 memory of the local accelerators. 45:11 And right now, we have laptops that-- 45:15 our MacBooks could have something in the order 45:18 of 100 gigabyte of memory. 45:20 And that means that we can fit very, very large model, 45:26 especially if we serve the model in quantized versions 45:30 like into it. 45:30 And so the larger some of the largest models 45:34 that are out there, we can take them and serve them locally, 45:37 which is a very interesting trend that 45:40 is following this other trend that a lot of user queries 45:44 are addressable by smaller models. 45:48 Next, please. 45:50 So the question that we wanted to answer here is that, 45:55 what role can local inference play 45:59 in redistributing the inference demand? 46:02 So we have all this traffic right now that pretty much 46:06 all of it is going to cloud. 46:08 Pretty much all of it is going to these accelerators like H100s 46:13 and TPUs and GP200 and GP300s. 46:17 But the trends right now are suggesting that maybe we can 46:21 do something different here. 46:23 And so let's see the next slide. 46:27 So in order to look into this more systematically, 46:31 we first define this new metric that 46:36 looks into not only the capability 46:39 or the way the models are-- 46:41 the accuracy of the models, but also on the efficiency side. 46:45 So on the capability side, we are looking this metric 46:49 that I'm about to define called intelligence per watt, 46:53 is looking into the percentage of the queries 46:59 that are addressable by the model for the single turn 47:05 and reasoning type of queries. 47:07 And when I say local models right now, 47:10 at least in this study, we considered 47:12 those are the models that have 20 billion active parameters 47:17 or less. 47:19 And on the efficiency side is how much useful compute 47:23 we can get from these per watt from running these models 47:27 on our local hardware. 47:30 So in short, the definition of intelligence per watt 47:33 is the average task accuracy divided by the average power 47:38 draw to solve this task by the model. 47:43 Next, please. 47:46 And we looked into a variety of models 47:49 hardware workflow type of tasks and evaluation metrics. 47:54 We looked into more than 20 local models like Qwen, GPT-OSS, 48:00 Gemma3 and so on. 48:02 We looked into both enterprise accelerators 48:05 and local accelerators. 48:08 The type of workloads we looked into, as you can see here, 48:11 there were 1 million queries from source from ChatGPT 48:16 and other reasoning benchmarks such as natural reasoning, MLU 48:21 pro and super GPQA. 48:24 And we also looked into the evaluation metrics 48:27 such as accuracy, energy latency, 48:30 the compute used and so on. 48:32 And all of these data-- so this massive data, 48:35 we used it to understand the trends, 48:37 but we are also open-sourcing all 48:40 of these kind of different metrics 48:42 across different type of underlying model and hardware 48:48 sub-trees. 48:49 Next, please. 48:51 So here is our findings. 48:54 It turns out that local models not only are very, very good 49:01 already, but the trend of their improvement 49:03 is also very interesting. 49:05 So since 2023, there was a 3.1x improvement in the accuracy 49:13 or the portion of the chat queries that they could solve. 49:19 This is very, very fast. 49:21 And this happened in only two years. 49:24 And right now, among the queries that I described 49:27 in the previous slide, they could address something 49:30 in the order of 88.7% of all these queries. 49:35 This is a very, very large number. 49:38 Another observation, which is kind of not very surprising, 49:43 is that local accelerators, in terms of their efficiency, 49:47 they lag behind enterprise chips. 49:50 For example, an Apple M4 Max, it delivers 1 and 1/2x lower 49:59 intelligence per watt than the B200. 50:03 And the reason for that is that these B200s are extremely 50:06 optimized to run language model workloads, whereas an Apple 50:12 M4 of course, is optimized to run these AI metrics, 50:14 but there are other kind of workloads 50:18 that are out there that they're optimized for. 50:21 And in general, while designing these chips, 50:24 the understanding wasn't that the LLMs are 50:29 going to be running locally. 50:31 That's why the majority of focus of chip designers 50:35 are on this cloud scale chips for running these workloads. 50:40 And lastly, there's this other very important observation that 50:46 the intelligence efficiency, or the improvement in this metric 50:51 is something in the order of 5.3x over the last two years. 50:55 So 3.1x of it is coming from better models. 50:59 Another 1.7x of it is coming from the improvement 51:03 in the hardware and how the hardware is becoming more 51:06 and more efficient. 51:07 So both of these-- 51:09 so both of these trends, one, is that local models 51:15 are becoming better and better. 51:17 They can solve problems that they couldn't solve before 51:22 at this very accelerated rate. 51:25 At the same time, model efficiency is becoming better 51:29 and hardware efficiency is becoming better. 51:32 This suggests that we are heading towards this future, 51:35 that more and more of this traffic 51:38 can be addressed, or can be solved by models that we can run 51:42 on our edge device, for example, on our laptop 51:45 or on our phone device in the future. 51:49 Next, please. 51:53 So let's go the next slide. 51:58 So with that, I get back to the IPW and the future directions 52:04 for that in a second. 52:06 But given this observation and everything else 52:09 that we have learned in this class, 52:12 there are a few directions that remain open. 52:17 There are a lot of interesting research questions around it 52:20 that are not addressed yet, and it's 52:23 going to be interesting to work on. 52:25 One is the foundational principles in test time scaling 52:31 and learning from this like new kind 52:34 of era of synthetic data and synthetic data 52:37 flywheel that these models are creating. 52:40 Right now, the way we are approaching it 52:43 is through RL, through collecting this data 52:45 and then fine-tuning the models. 52:48 But it kind of like suggests that there is more here, 52:51 our understanding of, first of all, why are we 52:58 seeing this kind of property? 53:01 What does it suggest from the model 53:02 that as we ask the model a question over and over again 53:06 through test time scaling, what is 53:08 happening that these correct answers are coming out? 53:11 What are the best practices to distill 53:15 this successful trajectories back to the model? 53:19 These are still open questions. 53:21 And there is a very clear segue from this synthetic data 53:25 flywheel kind of phenomenon, back to continual learning. 53:31 Something that seems like maybe it's not really we 53:35 haven't gotten it right yet, is the following-- like us 53:39 humans, as we solve tasks, as we solve problems and study and do 53:46 new things, there's this continual kind of progress 53:49 in how our brain develops and how we 53:52 become more and more skillful. 53:54 Whereas for the models, it seems like mostly there's 53:57 this offline process of here is like some agent take up 54:03 kind of experiences that are generated. 54:05 And then maybe after some time, there's 54:08 this fine-tuning process of the model. 54:12 It is not something that happens on the go. 54:16 And there is this kind of a mismatch between human behavior 54:21 and model behavior that a concept like continual learning 54:25 could potentially add answer. 54:27 Like, what are the new practices that we 54:30 can bring in model development that we 54:33 can bring these positive experiences 54:38 and learning from negative experiences and problem solving 54:42 that the model do back into the model in a more natural way 54:47 that is different from the current asynchronous, 54:51 like data generation and fine-tuning paradigm? 54:57 The last piece is the infra for the high throughput, 55:02 low latency test-time scaling. 55:04 So again, the way we do test-time scaling, 55:10 but there is repeated sampling, whether it's 55:13 these kind of back and forth in terms of updates 55:18 that we do to the previous generations. 55:20 We do tool calling. 55:21 We do this and that. 55:22 That is very different from the current mainstream chatbot 55:27 usage, which is just mostly single turn 55:31 back and forth with the model. 55:33 And what that means is that there 55:35 are a lot of opportunities for doing systems 55:39 and inference optimization work for these type of test scaling. 55:46 My lab did some of this work, like works 55:48 like hydrogen token SRS. 55:51 And these things are going to matter again 55:54 a lot more in the future because these methods are becoming 55:57 more and more mainstream. 55:59 So that means we need to have specific ways to deal 56:03 or optimize the systems and the underlying compute for them. 56:07 So overall, it seems like pre-training-- 56:11 looking at this figure on the bottom of the slide, 56:14 pre-training a lot of interesting work 56:16 in their discourse or our series of lectures 56:20 touched less on that, even though Akanksha 56:22 is an expert on pre-training. 56:25 But we mostly focused on post training 56:28 and the test-time scaling methods. 56:31 And there is this new kind of unleashed era 56:35 of like synthetic data, flywheel, and continual learning 56:39 that is happening in this kind of connection 56:42 between the fine-tuning and online learning 56:45 and test-time scaling that we hope that all of you 56:48 learned a ton about it. 56:50 But there's also a lot of interesting unsolved challenges 56:52 that you can address going forward. 56:55 Next slide, please. 57:00 Going back to the IPW metric and the observation about the shift, 57:05 the possible shift from everything 57:07 on cloud to a lot more locally also suggests new direction 57:14 in terms of inference serving engines that are hybrid, 57:19 that we can smoothly route traffic between our local 57:25 and cloud kind of models and accelerators, 57:29 depending on the need and the complexity of the resources. 57:36 The other direction here, again, is new model architectures 57:40 and kernels that we can use for energy efficient inference, 57:44 especially on the local accelerators that we have, 57:47 which this is like area that is much less kind of focused on 57:54 compared to the cloud accelerators. 57:58 And the other important piece is that energy 58:01 is going to be the most kind of valuable resource 58:05 that we have going forward. 58:07 And metrics such as intelligence per watt 58:11 and better ways of our understanding of measuring 58:15 energy and watt and power usage and optimizing for that, again, 58:20 is going to be very, very important. 58:22 And that is an area that is less worked on right now 58:29 in terms of the target metric of optimization. 58:33 But we expect that it becomes more mainstream and more popular 58:38 going forward. 58:41 And I think that's the last slide on the future directions. 58:52 Akanksha, you are muted. 58:58 There are some things in chat so. 59:03 I'll start with the first question here. 59:05 Do you think this is more pertaining to memory systems 59:11 or the alarms themselves? 59:15 I'm guessing-- so whoever asked this question, can you ask it 59:18 real time perhaps to put a little bit more context there. 59:22 Otherwise we're guessing the question. 59:25 Yeah, happy to. 59:28 Yeah, I was just referring to a few slides 59:31 back when you were talking about that continual learning 59:34 from failures and success. 59:38 Do you think that we're going to see 59:40 more of that in long-term memory systems and architectures 59:44 or what's designed around an LLM or the models themselves, 59:51 in terms of what's improved and what's worked on? 59:54 So I think the continual learning idea 59:55 that might pertain to long term-memory systems, that would 1:00:00 be the human analog of it. 1:00:02 But if you're trying to-- but before we even go there, 1:00:08 even in multi-step reasoning, learning from success 1:00:10 or failures or getting the model to keep that skill 1:00:13 set and update things in the weights, or in subset of weights 1:00:18 in a useful way, can you learn a new skill? 1:00:23 Learn how to learn is an important capability 1:00:26 that we basically don't have right now. 1:00:30 If you can watch videos, robots, try this. 1:00:33 When can you watch videos of how to do something and learn 1:00:36 how to do it, as opposed to being given a lot of task 1:00:40 demonstrations. 1:00:45 Azalea, do you have more to add there? 1:00:49 You're muted, if you're-- 1:00:56 Yeah, I think there are other ways also 1:00:59 to bring this knowledge, like obviously in context 1:01:03 learning is one way to enable continual learning. 1:01:07 One way I think about is imagine we had an infinite context 1:01:12 that the model could perfectly have access to every piece of it 1:01:17 and could learn from everything in it. 1:01:19 We don't have that. 1:01:20 But if we had that, maybe that was one solution 1:01:23 to continual learning. 1:01:26 Because we could put everything positive and negative 1:01:29 in the context, and the model could just 1:01:31 remember all of that at the same time and reason over all 1:01:34 of that at the same time. 1:01:35 But we don't have that. 1:01:37 And we go through right now, even 1:01:39 a million or a few million models capability 1:01:44 to reason over its EICL or in in-context kind of data 1:01:50 diminishes. 1:01:52 And there are other ways that we learned about cartridges 1:01:55 in the span of our lectures. 1:01:59 So that's one other way that we are enabling this long context 1:02:04 and in context learning without changing 1:02:07 the weights of the model, without fine-tuning the model, 1:02:11 but instead bringing them into the activations, or in this case 1:02:15 into the KB caches of the model. 1:02:18 So all I'm saying here is that there are other ways 1:02:22 to think about continual learning 1:02:28 without model fine-tuning, one could 1:02:30 be by just increasing the effective context 1:02:33 length massively. 1:02:37 And that could be another approach to long-term memory 1:02:43 and continual learning. 1:02:48 As a follow-up to that, in practice, 1:02:51 like in implementation, do you think, 1:02:54 what would be easier to do? 1:02:57 Constantly, updating the model's weights 1:02:58 or updating a memory store? 1:03:03 I think it depends on the application. 1:03:05 I mean, of course, you can argue that updating the memory store 1:03:10 would be easier. 1:03:13 If it's a knowledge base, if what you're-- 1:03:15 I mean, the simplest, my version of that 1:03:18 is that if I could keep a database that the LLM could 1:03:20 learn to look at, then I should just go update the database. 1:03:24 But what you're really trying to teach the LLM 1:03:26 is to the ability to reason over new domains. 1:03:30 And oftentimes having a side memory system 1:03:32 doesn't quite achieve that. 1:03:34 So that's where updating the weights does the job better. 1:03:37 And the example that I was giving in robotics 1:03:40 actually is very pertinent there as to no matter 1:03:44 how much memory systems you add, the robot which has learned-- 1:03:50 I mean, the cross embodiment generalization 1:03:51 that you were looking at on the lecture on Monday, 1:03:54 for example, that doesn't happen if you don't update the weights 1:03:58 just by having memory systems. 1:04:00 So that's more of a skills transfer problem. 1:04:04 Makes sense, yeah. 1:04:05 Thanks so much. 1:04:09 There's one more question, I think from-- 1:04:12 what he's saying is that in the absolute zero paper, 1:04:17 the environment seems like the data used for post-training. 1:04:19 Is there a paper for agents to self create environments? 1:04:25 I mean, if you go back to the worlds 1:04:27 where there was narrow intelligence, 1:04:31 there you could basically create simulations for games. 1:04:35 And those were used as environments for toy tasks. 1:04:41 The reason environments matter in the current generation 1:04:44 are that they are proxies for real-world tasks. 1:04:47 So it's not so much that there's a paper for agents 1:04:49 to self-create environments. 1:04:51 It's more along the lines of what is the set of tasks 1:04:53 that you're trying to represent. 1:04:55 And if there is an easy way to simulate them, 1:04:58 then whether you use agents to create 1:05:01 that or software to create that, that's fairly straightforward. 1:05:04 But is it a reasonable proxy of how 1:05:07 the model will interact with the real world to get feedback. 1:05:10 And for gaming, it's kind of a finite space to explore. 1:05:17 So it's easier to represent that in code and have simulations. 1:05:23 OK, thank you. 1:05:26 More questions from the class? 1:05:41 You were saying something? 1:05:42 I think up here is a question. 1:05:46 We can't hear you because you have to speak up. 1:05:48 I think that there are no further questions. 1:05:52 OK, awesome. 1:05:53 OK, cool. 1:05:55 OK, so let's start by-- let's end by thanking the class. 1:05:59 It's been a real pleasure teaching you all 1:06:03 and creating the content for this class 1:06:05 so that you can learn about the latest and greatest set 1:06:08 of techniques in the self-improving agents area. 1:06:12 It's a evolving area, so anything that we teach now 1:06:15 starts to be history by the next time the class rolls around. 1:06:20 And yet the basic techniques that you're learning 1:06:24 are extremely valuable over time, 1:06:27 because a lot of the new stuff still builds 1:06:30 on top of the basic techniques. 1:06:33 So really grateful to have worked with you all 1:06:36 and looking forward to what you do 1:06:37 in your projects and your posters. 1:06:40 I'll let Azalea also thank the class. 1:06:43 Yes, thank you so much for-- 1:06:49 the way we created this class is that we were also 1:06:52 learning about a lot of things as we were creating the slides 1:06:57 and we were preparing the course material, 1:06:58 because a lot of these topics are just 1:07:00 like so fresh and so new. 1:07:03 And we were excited about them and we wanted 1:07:05 you to also be aware of them. 1:07:07 So thanks for accompanying us in this journey. 1:07:12 And we are very grateful to you and we 1:07:14 hope that you learned some new skills, 1:07:20 you got inspired in some new directions, 1:07:23 and we really hope that this is just the beginning for you 1:07:28 to go ahead and do so many more amazing work in your research 1:07:31 projects, in your jobs, and in the future in general. 1:07:34 So thank you so much.