As far as I know, there isn’t an interface like nvlink that allows these macs to work in tandem; they would just send their data over ethernet, maybe thunderbolt/usb-c?
Nice, not sure why I was downvoted for suggesting the cheapest GPU which comes with quite a lot of VRAM... but I guess some people are fatigued by AI / GPU talk?
From what I gather, the Chinese are behind, but a lot of their research amounts to scrappy, clever discoveries in how to use more novel technologies (for Qwen and Deepseek, its mixture of expert models, that can do inference using a portion of the model at a time). The chinese also distill information from American models, so there’s that.
The American companies, from my impression don’t involve themselves with such lowly “hacks” because they have so much money to just push forward with doing everything on big heavy models that run on the most cutting edge nvidia chips that they can, the moment, kinda sorta get on demand (I say that in some degree of jest).
this is not an effective long term strategy in a collaborative environment that is advancing for the same reason that having a private secret fork of the linux kernel with a few proprietary improvements is not an effective strategy.
integrating your own work with the latest public advances takes resources. For one or two small changes this is manageable, but the further you diverge from the public, the cost of maintenance rises exponentially if you want to continue to integrate public advances. when you publish your meaningful advance, you offload the maintenance burden onto everyone else (and they only have to pay a linear cost rather than an exponential one) as it's integrated by default in new work.
In most cases, the (exponential) maintenance cost of integrating public advances with secret ones exceeds the value of the public advances, so most that undertake this strategy of advancing the open frontier in secret don't attempt to integrate continually, but instead try to make a breakaway sprint in isolation to grab a few sticky customers before the unstoppable wave of the public frontier catches up.
This is a pattern commonly seen in university research departments when researchers switch into product development mode, most of these projects are a sprint to advance away from the public frontier once a good idea is found and they do good work and find a few customers for a little while. But if you check back in a few years you won't find an advanced research department but a zombie IP company that brings in a steady income via IP enforcement and a small number of customers for whom switching is too expensive.
The concept of open source doesn't really apply to AI models since their behavior is mostly controlled by the data they were trained on and the complex ways they are trained. Having the source code of the model by itself wouldn't help you.
From a practical POV having all the training data, training infrastructure, and training know-how wouldn't help you either unless you could afford to spend the millions of dollars (hundreds of millions for a SOTA model) in compute to train it each time they released a new training set, in which case you're only talking about the big commercial companies. "open source for the people" just does not apply.
Publishing RL/SFT/self-distillation harnesses would be very impactful even without the data.
Particularly when it comes to tool use w/ self-distillation it can be done without any data... have a tool the model doesn't know? a teacher model RTFMs and the source code, and helps the student learn to get it right.
> If (and that is a big if) the concept of open source doesn't apply, then the term shouldn't be coopted to mean something else though.
Yes, but for whatever reason this usage seems to have stuck. Open weights is definitely a better name. I assume the reason "open source" has stuck is because you can download and use it for free, but "open source" was always intended to be about "free as in speech", not "free as in beer". That said, I remember when the term "open source" was invented, and it was always a bit different, more commercially aligned, than the goals of the FSF.
> But even if I can't build it from source locally, being able to see what went into the model is an important part of what open source is about.
True. Unfortunately LLMs have become such a big money and closed enterprise (the opposite of OpenAI and Anthropic's altruistic founding principles) that it's hard to see these commercial models releasing their training data, especially since this data is the closest thing they have to a moat other than the cost of training.
The most valuable training data right now seems to be "reasoning data", and the need for this at least may disappear as AI moves beyond pre-trained language models to smarter systems capable of learning for themselves, and that can actually reason, not need to parrot reasoning data.
When people in NYC are driven out of their neighborhoods because of gentrification, they generally move down south. There isn’t some magical part of town that they can afford with their “current budget”
Economic theory says some things are theoretically impossible, no literally, but economic theory wouldn't say that here:
The local housing market is much more complex than supply and demand, with larger economic factors (e.g., interest rates), very imperfect information (affecting everyone from buyers, to sellers, real estate agents, lenders, etc.), coordination by landlords (e.g., RealPage), non-economic factors such as prejudice (or just a co-op board!), government actions, larger trends, temporary inefficiencies, etc.
Economic theory is useful, but it does not predict or circumscribe the immediate reality of individuals. Life is much more complicated than that.
First, I didn't say there is no supply effect; I said it's far from impossible for the effect to make a difference.
Second, many factors are involved in a complex market; you and I don't know how much effect the supply had in this case. That you are interested in that input isn't evidence of its effect.
Any program created by the US government can be captured and handicapped, like has been done to the USPS.
Also, the Postmaster General was on Capitol Hill today saying how this time next year the service won’t be able to afford delivering to all addresses in the US.
> Any program created by the US government can be captured and handicapped, like has been done to the USPS.
Agreed but even despite that they generally are a net positive.
> Also, the Postmaster General was on Capitol Hill today saying how this time next year the service won’t be able to afford delivering to all addresses in the US.
The same postmaster general who is a longstanding board member at FedEx.
And the US Post was still an extremely effective agency for well over 150 years, only truly beginning to become shackled when Nixon transformed it into the USPS in the 70s, and even then it retained most of its efficacy until the 2000s and 2010s when it truly began to fall onto its last legs.
But also despite being shackled the way it currently is, it's not exactly nontrivial to reform it provided there was any political momentum towards doing so. So it may get its legs back in the days following this administration.
You can connect keyboard, mouse, and video, but you just get screen mirroring on the screen (or it can properly display a full-screen video in some apps, I think... though DRM video may refuse), so, it's pretty limiting.
Sure, it can flow down to regular people now, because regular people don’t have access to 10s of millions of dollars to train a trillion parameter llm…
reply