Machine learning objectively is you have some data, you create a mathematical structure which will be trained on that data and soon give you some output. The data can be actual or synthetic, and that mathematical structure. Can just learn things if you let it run properly, and sometimes the data might be so garbage yet it works, or the data is so. Strange, it gives you results for things you might never expect. We see this in trading all the time. Weird is form of alternative data giving you prediction power in stock market. So that is high strangeness that synthetic, unrelated or weird data giving you outputs related to something very valuable. Valuable just because you created the mathematical infrastructure around it.
And I think with LLMs we might be in an era where it’s not only mathematical structure that can do it, and it’s not only numeric data that can be used to create these systems. We are now in a time where you can have actual image or text or internet scraped data. Create prompts and get output and keep iterating on that prompt autonomously until you get the output you want. Or you can just go for synthetic data or fake data. Or sometimes maybe you can create a reference model of your own that okay this is how I would make it. These are the prompts.
Now bitter lesson me, motherfucker. Go do to me what you did to chess, go and software engineers.
Create your own type of data, your own prompts that will deal with that, and show me you are making output better than me. Maybe I’m a human, maybe I’m dumb, maybe I don’t know the right type of data or the most optimized type of data. You go, create a fake one, create prompts accordingly, and go fucking do it. Go do it better than me. We are now supposed to ask. For solutions that can be way novel and better, and stranger and beyond our understanding, because these guys have crossed the thresholds a month back that is just highly strange. So what I’m essentially saying, ask your LLMs to create. fake or synthetic data. Ask it to create what are fake prompts it wants to create so that you have a infrastructure around your process and ask it to run the process of applying that prompt on that data and what are output comes. You should give it clear target that okay this output is supposed to be like that. Let’s see if that works.
objective
→ model invents useful data
→ model invents useful representations
→ model invents useful prompts/procedures
→ model tests them
→ model generates adversarial examples
→ model revises its procedure
→ output
→ evaluator compares output against target
→ repeat.
“Determine whether prompts are even the correct abstraction. Determine whether my reference examples are useful. Invent better examples if necessary. Invent synthetic edge cases I would never think of. Invent intermediate variables that make the task easier. Run competing approaches. Preserve whatever reliably moves the metric.”
“LLM, I’m not going to tell you how to do market research / design / coding / video ideation / valuation / copywriting / whatever.
Here is the environment.
Here are examples.
Here is what counts as a win.
You are allowed to manufacture your own training material and procedural scaffolding.
Now bitter-lesson me.”
I think we gotta stop being the bottlenecks inside the weird training situation we have got going on with LLMs, where we use our innate knowledge so that it can understand what we are trying to do. I think you should just show up with examples of what you want and what is a clear win condition. That’s that’s all we need.It can prepare its data. It can prepare its prompts. It can prepare its process. It can bitter lesson you hundred percent.
I think what this does for me is that now I’m not even thinking about the solutions or systems. I just have to figure out what I want and tell you that hey, go figure it out, go high strange high strangeness it out. First, me writing blogs to understand marketing was the bottleneck because how many can I read? Then I switched up to learning from the legends that going through all the niche blogs and really really good books is how I accelerated. So instead of information self discovery, I went to looking for. existing knowledge bases and now I’m like okay even going after those maybe I might pick a wrong book for my system so why not just let high strangeness do it all for me I just have to say or push around what I want
The bottleneck is now just me understanding what I want. Damn, we are so far ahead.