Mostly useless reading list. Very little emphasis on technical SOTA and mostly policy level waffling.
And regarding the data question the other commenters are asking — you scrape everything you can (oh look I used an em dash, wanna run me through the cover-your-ass Pangram?). Anna's archive, The Pile, the various Huggingface data sets and aggregates, Common Crawl. You pay proxy farms like Bright Data to run residential and mobile gray area proxies and VPN and CloudFlare bypasses to do more scraping. I see a lot of HN users up in arms on the front page thread about LG TVs having VPN SDKs within them etc. And then five minutes later they will go back to their frontier LLMs to print their pay cheque. Fucking hypocrites of the highest order. It's oh noes how bad this tech is in my house but we will happily benefit from it.
Then you have a data cleaning team deduplicate and clean and annotate the data (with or without help of more AI)
Certain RL specific datasets for supervised fine-tuning and RLHF like coding and git commits and chat needs to be curated by hand depending on your use case.
The level of discourse on AI has fallen tremendously on HN if 5 years into the AI revolution people are still wondering why datasets aren't being released. They aren't being released because they are a fucking snapshot of the internet for fuck's sake. There are a few "sanctuary" nations where AI data scraping is somewhat legally unenforced but the United States is not one of them so stop asking why data isn't released on a Bay Area website. Use your head for once. Too many React and YouTube influencers and the brains have been rotted.
Out of all the so called "AI engineers" here pontificating about "alignment" and "AI safety" and "Recursive Self Improvement", I wonder how many can even formulate or describe what an ELBO is. I wonder how many product managers here yapping about "recalibrating their priors" actually know what a prior is. I truly wonder why LLMs seem so magical to people when it's only a few steps removed from the same neural networks people have been using since 2015, at least architecturally (except scaled up by a few magnitudes).
Get your head out of Roko's Basilisk's agentic ass and maybe actually read the technical reports and papers for once.
(One important paper post Attention is All You Need is the DeepSeek paper where they used RL to bootstrap the "thinking" chain of thought token chains. IIRC it's the DeepSeek R2 paper. That's one of the most important papers for understanding LLMs beyond basic ML neural networks. If you need a quick way to get up to speed, read that one).
> I see a lot of HN users up in arms on the front page thread about LG TVs having VPN SDKs within them etc. And then five minutes later they will go back to their frontier LLMs to print their pay cheque. Fucking hypocrites of the highest order. It's oh noes how bad this tech is in my house but we will happily benefit from it.
Something something participate in society
(Or more elaborate: Disagreeing with the status quo does not make you a hypocrite for benefiting from the status quo.)
I generally take the technocratic view of things, in other words you should understand the things you are trying to ban/regulate .
Replace AI with vaccines and the problems with how some parties are pushing for regulatory capture and/or blanket bans become obvious. Build policy around the effects e.g. job loss, biases etc. but go full laissiez faire and hands on on the technology. Nukes are 2nd amendment.
On the technical side, there's a lot of grunt work that's completely unrelated to machine learning but underpins modern AI. E.g. high performance computing and numerical techniques have zero relevance to day to day LLM research but makes or breaks the implementation. There are better reading list for those but you need to at least have a vague understanding of what numerical methods or a math kernel is before starting to lecture others on the economics of ML scaling and hardware. For stuff like data cleaning and scraping, it's a well known gray area field so you aren't gonna find too many for dummies guide on it. Can't have honest discussions on hacker news either because the people here get their panties in a twist over data scraping despite many of them doing it with zero hesitation or compunction if it comes up in a jira ticket. Proxy farms are the unsung hero of LLM engineering.
The Boeing 767 does not have a Flaps 40 setting. The "Forty" callout was an automated annunciation of the radar-altimeter measured altitude above the runway.
Most web devs have never heard of Colo unfortunately, they only know Vercel and AWS. Hurricane electric should sponsor more booths at colleges. If they give out more swag maybe the millenials and gen Zs would finally understand bandwidth pricing.
My friend's vibe coded vercel site got hit by Meta for 21 million page views in 2 days. It cost him over $300.
My self hosted compose stack running in my basement with two 9's of uptime was a 1 time cost of $600 between cat6e, refurb mini PCs and tons of time prompting for NixOS flakes that met my needs. I'm not sure who came out ahead.
> The AI thus builds confidence over time, acting on object detections that persist across several frames, rather than, say, slamming the brakes due to a camera blip on a single frame.
Nice in theory but in practice they still need a ton of training data. The bitter lesson is very bitter. Tesla's end to end FSD uses occupancy networks and brake stabbing still hasn't been fully solved.
While cc banks are in these times starting to side more with vendors on charge disputes, its still far more recourse than bank drafts offer you in a transactional dispute
And regarding the data question the other commenters are asking — you scrape everything you can (oh look I used an em dash, wanna run me through the cover-your-ass Pangram?). Anna's archive, The Pile, the various Huggingface data sets and aggregates, Common Crawl. You pay proxy farms like Bright Data to run residential and mobile gray area proxies and VPN and CloudFlare bypasses to do more scraping. I see a lot of HN users up in arms on the front page thread about LG TVs having VPN SDKs within them etc. And then five minutes later they will go back to their frontier LLMs to print their pay cheque. Fucking hypocrites of the highest order. It's oh noes how bad this tech is in my house but we will happily benefit from it.
Then you have a data cleaning team deduplicate and clean and annotate the data (with or without help of more AI)
Certain RL specific datasets for supervised fine-tuning and RLHF like coding and git commits and chat needs to be curated by hand depending on your use case.
The level of discourse on AI has fallen tremendously on HN if 5 years into the AI revolution people are still wondering why datasets aren't being released. They aren't being released because they are a fucking snapshot of the internet for fuck's sake. There are a few "sanctuary" nations where AI data scraping is somewhat legally unenforced but the United States is not one of them so stop asking why data isn't released on a Bay Area website. Use your head for once. Too many React and YouTube influencers and the brains have been rotted.
Out of all the so called "AI engineers" here pontificating about "alignment" and "AI safety" and "Recursive Self Improvement", I wonder how many can even formulate or describe what an ELBO is. I wonder how many product managers here yapping about "recalibrating their priors" actually know what a prior is. I truly wonder why LLMs seem so magical to people when it's only a few steps removed from the same neural networks people have been using since 2015, at least architecturally (except scaled up by a few magnitudes).
Get your head out of Roko's Basilisk's agentic ass and maybe actually read the technical reports and papers for once.
(One important paper post Attention is All You Need is the DeepSeek paper where they used RL to bootstrap the "thinking" chain of thought token chains. IIRC it's the DeepSeek R2 paper. That's one of the most important papers for understanding LLMs beyond basic ML neural networks. If you need a quick way to get up to speed, read that one).
reply