• 0 Posts
  • 599 Comments
Joined 2 years ago
cake
Cake day: August 26th, 2024

help-circle
  • I’m probably going to stumble over some of the terminology here, but I think it might be possible to describe what @[email protected] is proposing as a consequence of LLMs ultimately being lossy compression systems. Inference is a function over a lossily-compressed data set, and “chain-of-thought reasoning” and “agents” may sound sophisticated, but are simply applying containerization and DevOps tools to VM images of the inference application in an attempt to get around hard memory limits on the context window for inference. “Chain-of-thought” attempts this in a serial fashion, passing results from one instance to the next, while “agents” implement this hierarchically and recursively (and woe to the poor bastards who wished that mess upon themselves). But in both cases, the “finalization” phase is necessarily a further lossy compression step, attempting to compress a result from the inference process to a fresh instance of the inference application, so as not to immediately blow out the new instance’s context window.

    Given this necessity, it comes to seem somewhat intuitive that there may be “strange attractors” in the higher-dimensional vector space that is the compressed data set which surround code that creates and maintains message passing channels. No matter what you’re doing with an “agentic” process, the inherent necessity of context cramdown & message passing means that querying into the space where such code examples lie is a hidden requisite of running the damned things, thus turning such functionality into the sort of selfish elements that BioMan is talking about.

    The problem in investigating and concretely describing this phenomenon is nailing down the exact functions and processes that make it happen. Given the godawful messes in the Claude frontend codebase that @[email protected] has been documenting, I’d be surprised if there’s one developer in a hundred at Anthropic or OpenAI who can describe in detail how the intentionally-developed context-passing code for their “agents” works.















  • Yeah, I began losing interest in Greer as it became clear that he was perfectly happy squatting in the middle of the red-brown alliance during the Trump era. His critiques of industrialism and unquestioning belief in technological progress broadly align with what we discuss here, but he will always coddle MAHA types and tale a shrugging “well, what can ya do?” attitude towards people like Trump, as it fits his preference for cyclical theories of civilization.

    I noticed a couple months ago that he actually managed to dig Nick Land out of whatever tweaker den that guy’s been hiding in for a podcast, which says a lot about what he’s willing to indulge these days.






  • I am personally surprised the Chinese state continues to allow the open sourcing of models this good, given potential risks. To be clear, I myself might be fine with models presenting this level of marginal risk being open weight, but I am surprised that China is fine with it. I suspect the reason they are is 75% explained by strategic blindness/lack of AGI-pilledness (the CCP is very Yann Lecun-y in its views of AI).

    The strategy seems quite clear to me, and not blind at all: they want to hold a gun to the head of over-capitalized Americans like Mr. Ball and his enablers, and will be perfectly happy if that drives our tech industry into a financial escalation spiral that ultimately neuters it for years to come. He has the barest glimmerings of being aware of this, but frames it wrong because he can’t let go of the sci-fi sentimentality that drives his whole industry.

    The other 25% or so is their lack of compute for customer inference (making China’s open-weight strategy an unintended byproduct of US export controls) and the normal Chinese strategy of aggressive exports.

    The Chinese still have overwhelming clout in final electronics assembly, i.e. turning this shit into things people can actually use. Of course they are interested in maintaining that advantage. And if, in the medium term, they can nudge more financially desperate American politicians and executives towards loosening trade and investment restrictions, possibly even towards equity partnerships, so much the better. “Lack of compute for customer inference” is ultimately a red herring. If I’m a Chinese economic strategist, why would I not expect GPU tech to become commoditized over time, like every other aspect of advanced computing technology has? Perhaps even more rapidly, as the designs are inherently massively-parallel implementations of similar base execution blocks!

    The rest of his “accelerationist”/“decelerationist” drivel concerns a silly spat that is mainly centered around and driven by Twitter. It is still difficult to overestimate just how brain-rotted Twitter has made these people.

    e: case in point, he leads out with barely-veiled Covid lab-leak trutherism, surely just to rile up the hogs. What a loser hahahahahaha

    “A nonliving, invisible, dangerous, and infinitely self-replicating agent escaped from a Chinese lab,” you say? Color me shocked.