Hacker Newsnew | past | comments | ask | show | jobs | submit | Imustaskforhelp's commentslogin

> The fundamental problem is simpler: no one will slow down because no one trusts anyone else to slow down.

I would argue that nobody trusts anyone else in the case of AI/AI related stuff.

The amount of money involved in the space is astronomical and incentives are really misaligned. AI is such an extremely polarizing topic.

A meta example of this distrust is that I can't (really) trust Dario Amodei's comments itself! (is he saying this for more future IPO money or out of genuine fear) which was ironically also what your comment's about as well.

By following the same chain of events, we can have wildly polarizing claims about the same event and its explainations.

I seriously have to wonder what historians will have to say about this period of human history.


> The reality is, capabilities have largely converged across foundation models over the last 18 months, and much of the value add is coming from the harness layer itself now.

Can you please elaborate on what are your thoughts on open weights models (GLM 5.3, Kimi K3, deepseek etc.)

and if the value add is coming from the harness layer itself, then thoughts on open source harnesses (there are so many harnesses but to name a few: opencode, pi [omp as well], maki, codex is OSS as well, fx.sh) and you can always combine them with skills (Obra/superpowers, matt pocock skills plus using these skills and others to create some other custom skills tailored to your use case as well)

And what about the combination of both now with this cheap open weights models + open source harnesses and other things to compete over the closed garden ecosystems?

How does that comparison follow in reality

Could you in theory use these methods to save on the massively expensive $$$ token spending on Anthropic/OAI?

(Personal anecdote but I have GLM 5.3 + maki [sometimes omp/opencode but mostly maki] and its good enough for most use cases out there that I have and I dont know of too many use cases outside of say recreation of games for examples maybe that I would prefer complete SOTA models. I would also love to know where you believe that SOTA models absolutely do still make the difference discounting the benefits provided by the harness.)


> Look at image gen: you press the magic button and get a 'Pelican on a Bike' - but you can't just change the 'hat' of the Pelican. You have to press the magic button again, and you get a whole different Pelican on a Bike.

I do understand what you mean and perhaps I am trying to treat it as a problem to be solved and challenge accepted but my first thoughts are if its an vector image like SVG (the famous simonw's pelican on a bike benchmark)

Then you could in theory have a layered approach and then just change the code of the hat so for example (IIRC) <--Hat--> Code. <-- Bicycle-->

So basically I am suggesting modularizing of concerns of areas so that you could have better autonomy over what exact thing you wish to change.

Though I imagine that you and I might be saying the same thing and you are suggesting that the crux of the argument is exactly that you need specific know how in reaching to that said modularity where you can best use AI.

When you might need the generation of the SVG as compared to image generation because you realize that your project might need the pelican changing lots of hats and image gen wouldn't be feasible but for that to actually say to AI, You might need to know in the first place that SVG or (HTML?) might be better use cases for this. Again, I must admit that I am not an expert in this so I can't absolutely comment on what the best thing for this particular use case could be and maybe that might be your point.

Have we arrived at similar conclusion or perhaps, is there more nuance to it?

On a side note: I actually once had a real use case of where I needed layered approach similar to Figma but generated through AI. preferably something which can just work through good ol chat app UI which could give an index.html or other code for this modification/layered approach.

Like recently there is https://bento.page which has been somewhat similar for this in some sense but for pdf's. So is there something different but for image-ish thing? If there is an expert lurking here, I would love to know the answer and gain some knowledge about it, thanks :-D


Yes, for things we can construct, and where we can train on the pieces, there's hope - software is a bit like that. Its not like 'raw' image generation, it's not a perfect example but the notion remains.

On some occasions, as I have come to understand Hackernews, its not really about the content but sometimes just about the Title or literally just the sentiment behind it.

Maybe we are all being shown recreation of Minecraft or other video games as a benchmark and maybe there is a discussion to be had of the statement of (is it a benchmark or not). I, for one, just appreciate that we are all here who could talk about this topic that I was just thinking about and can have a good discussion about.

In some sense, its speaking some thought that everyone (many) might be thinking internally out loud.

Also on a related topic, I have a question that I wish to ask, there are many text files that I have which are me just writing extremely crude thoughts. I once tried to let AI re-write it in a blog format and I found it to somewhat express what I was saying, perhaps saving me hours.

I still don't upload those blog posts because I don't wish to write AI generated blog posts even though I might have spent more than hour writing those original thoughts. To me, its because I don't exactly know how to disclose things. Some of the original thoughts for example might be quite PII/sensitive. Should I share the exact prompts or the original thoughts and what are the correct hygienic ettiquetes towards this nuanced topic?

It would also take me more than hours to edit the posts and I might still not be confident in my ability of trying to create that blog post and might end up never finishing that public blog post in the first place due to procastination.

I also don't just want to sometimes write AI generated blog posts because its a slippery slope. Without proper disclosure, the other person might not even know if another human on another side of planet even gave a thought about it or not. For all intents they might think that nobody really wrote it but only an AI and just "write a blog post about X"

I also would dislike being the person on that end and I also dislike reading AI generated posts. It produces visceral reaction of hatred and closing the tab when I read AI generated slop.

I ironically wanted to create a blog post about this nuance but I couldn't capture it and I then had AI write that blog post but I didn't post it. I have had some very nuanced meta experiences about this exact topic to be honest and I would love to hear some thoughts & opinions about it.

I have the thoughts but because of all of this, in some sense, a lack of means to share. I would wish to hear other peoples thoughts about those thoughts that I had and if other people resonate with it or not but alas, some of them are just text files within my computer.


> What else is left, once all these benchmarks get saturated?

I was thinking this thing quite recently and wish to write a blog post about it but IMO, I believe that the next AI benchmarks would be profitable businessmaking. There are/were already some trials done (The famous WSJ/Anthropic one[0]) and some shop iirc in SF which is doing this and recent example of trying to start a lemonade stand with AI[1]

In my opinion, it can't really get more meta than that plus on the more humour side, it could help these companies make some dollars as they desperately need it.

I always wonder if these AI models can self-autonomously generate profit, why would these AI companies try to let you generate the main profit while you just pay them one time or a very small amount in tokens. Why should these AI giants allow a person to be a middle man in the first place in some sense and instead not just run these agents autonomously.

Maybe it could be because these AI agents might not be the best in such a meta task (money making/business) but as such, it does feel to me that it might be a good benchmark in the first place.

[0]: We Let AI Run a Vending Machine. It Lost All the Money. | WSJ: https://www.youtube.com/watch?v=SpPhm7S9vsQ

[1]: Everything AI Does When You Ask It to Start a Lemonade Stand: https://www.youtube.com/watch?v=6Ide5pRLR8Y


I would be surprised if we're not there yet. If Jane Street can generate $500 million per MW of compute right now I'm sure it's already possible.

We might still be figuring out how to benchmark these models by the time next gen comes.


I really like these tests for what its worth and I see them on youtube sometimes. I would like to ask a few things though

TLDR: Basically focusing on recreating pay to win (mobile or otherwise) games and recreating them non pay to win perhaps instead of focusing on recreation of minecraft for benchmarks could have a genuinely meaningful impact, and making these games portable as well could be another interesting idea. [so it can be played on any operating system/device so using web or if native then for (Android/IOS/Linux/Windows/MacOS) using game engines like (Preferably godot)/Unity/UE.]

Could the test focus more on pay to win games with unique dynamics.

For example: I literally wanted to create a clash royale recreation because clash royale is a highly pay to win game.

The game is unique enough to have memories but is pay to win enough that it ragebaits me as to what its current situation is, its so pay to win now. A recreation would have genuine effect whereas yet another minecraft clone wouldn't.

I have some fond memories of the game and my brother and I used to play it (my brother moreso than me). Also clash of clans and clash royale famously prevented windows users. So I remember downloading bluestacks to play it on laptop but it required 2GB of ram and back then we only had 1GB. (Ironic that we might come back to that time)

Another question that I have for you which I have been genuinely curious is: who is footing the bill for these benchmarks and youtube videos. What are the economics surrounding it?

I imagine the bill to run quite hot sometimes and I find running these benchmarks to be quite unaffordable personally.

I also wish to ask if you have any theories as to why not people on Youtube share their videos. I found this [0] Minecraft clone by Fable 5.1 extremely good yet they haven't shared the source. I am unsure as to what exact reasons might be behind most youtube videos on recreation with AI to not share the actual code. I don't find much rationale in not sharing AI generated code of a recreation of a game especially if one is making a video about it. So thanks for once for actually sharing the output code as well as I surprisingly found it to be a bit rare!

[0]: https://youtu.be/I0do_vbnMBI?t=361


I'd make a distinction between youtubers who are incentivized to dial their reactions to the max for everything[0] and people doing silly tests but keeping their expectations and reactions real (most famously, Simon's pelican test).

As to why the prompts aren't shared, for these more complex things it's most likely a somewhat messy process (ie. not a single-prompt one-shot creation; some back & forth) that would make the whole thing seem less spectacular.

Likewise for the end result - it's probably cherry-picked what works well. For example, in Claude models' resuts I always get stuck in water (something about height/jump calculations is off), where with Astra I didn't have that problem. These sorts of issues you can only spot if you try to playthrough yourself.

This is just my speculation tho. As for me, I do the tests because I'm interested in the results (easy comparison across time & models) - then I started sharing them because people asked.

I have OpenAI and Anthropic subscriptions so testing these is not an extra cost for me (I'm sloppy around recording the tokens & API-equivalent cost tho - have to improve on this). For the other models, it's total a few bucks per month or so.

So if you're careful about the cost, it's not too much, especially for a serious youtuber who's doing it for commercial reasons.

Finally regarding your comment about cloning the popular enshittified games - I don't think people are going to be doing that for testing, but if you want to have a different spin (or do a close-enough clone for yourself) on a game you loved, the modern AI systems can often deliver!

[0] from the link you posted "i am in disbelief, this is insane, bro what is this, you can't believe it's ai" - yeah...umm, it's not that good :)


> Thanks. Its insane how "Twitter, but for tech nerds" has a hard requirement to either enable Javascript or download some app.

https://masto.mirror.forum/beige.party/@intransitivelie/1170...

Source code: https://github.com/SerJaimeLannister/mastoview

(Disclaimer: It's vibe-coded. It does server side rendering to then just give pure HTML to the end user with no JS required.)

I hope that this helps people who want to view Mastodon without JS.


Offtopic but I remembered one HN comment about where someone was working for an internal app for a company for the fast food/ice-cream or food app in general (or maybe it was for the contractor)

I think some people found that job really peaceful when others might've found it more frustrating.

To me, I think though the issue isn't the scale or even the internal part of it but rather in this case, there's a difference because its being used infrequently (1 hour) as compared to that other HN comment was still being used quite frequently (it was used every day)

Though, I must admit, I don't find it too frustrating. I can be wrong I usually am but, Sysadmin/DevOps/Security engineer does the work as well where sometimes 1 hour can matter much more than months and they do work for months so that they don't get to face that 1 hour issue just because of how devastating it can get, Though Sysadmins do a lot of regular work as well.

There are many jobs where they exist for reasons where the employee exists to handle the bad things. You would wish for the bad thing to not happen and the employee might feel like they might not be doing much if the bad thing doesn't happen, BUT when that bad thing happens, You would be happy that they would then be there for you.

Maybe not a direct 1:1 comparison but hopefully I can express my point. I feel like there are many jobs in tech, or maybe in the world in general which fall into a similar fashion?


I would agree with you on AI/LLM being more than agentic coding but at the same time, I think there's more nuance.

For example, PDF's and powerpoints can be generated using agentic coding by things like https://bento.page or other ways of generating them in an agentic coding fashion.

A lot of browser automation could/is also done by agentic coding.

It can also help them set up and configure self hosted software with the help of LLM's and debugging if its working or not.

You can create videos using Manim and remotion.dev and also excalidraw-animate and generate excalidraw files agentically if what you need is more vector style graphics (which surprisingly can fit into many ideas) rather than say a real life human waving video/more photo-realistic video (but I must say that this has certainly its own pros/use-cases as well).

It might sound self-explainatory but turns out that coding can represent a wide range of problems!


I get that. I use agents a lot and LLMs often reason with code. It is valuable. I just think the floor is a lot lower for general reasoning and common tasks like that. And in 6-12 months it won’t matter. Google will publish better models. The temporal distortion of how long a Sol or a Fable has existed is real. No one is suddenly missing out on some giant competitive edge because their model is a few months behind. I feel like it’s all just going to normalize and things other than how well your model can write code will matter more and more in 12 to 24 months.


Sure I understand what you mean as well and I am not asking for SoTA models to be created by Google but more so explaining why coding is still the largest focus for many labs.

I personally wish to get more smaller models (like the recent qwen model) and other open source models like GLM 5.3 and the glm flash model.

> No one is suddenly missing out on some giant competitive edge because their model is a few months behind

Sure I can agree with that. The competitive edge might still exist but I do get the underlying sense of what you are trying to suggest.

> things other than how well your model can write code will matter more and more in 12 to 24 months.

What are the things then which you feel like could be more differentiative factor? For example, I personally think multi modal is still quite preferrable in AI models. I use GLM 5.2 and it doesn't have vision and I can certainly imagine time/use-cases where multi-modality would've helped coding and even other use cases as well. So what are some other use cases that you are thinking? Video generation models like Veo/Sora?


/me points over at Flash 3.8 :)

> To overcome the relatively limited compute and memory capacity of individual chips, we built a dedicated inference engine for this architecture on top of SGLang. Notably, this effort was accelerated by our GLM-5.3-powered infrastructure agent, which assisted engineers in developing and optimizing kernels, diagnosing performance bottlenecks, and improving the serving stack — creating a feedback loop in which the model helped optimize the system serving the model itself.

> (...) Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale.

It might be one of the most actually practical tasks that AI might've done because the compounding effects of it and also its implications are/feels so immense. It feels as if Nvidia might be in a slight turbulence from it.


Now translate this to physical world, robots building and optimizing other robots... getting iRobot (2004) vibes


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: