> The phrase "frontier model" is starting to mean two things. One is a checkpoint. The other is a system boundary.
LLM-isms aside, I don't think we want this to be the case? An LLM, for all its complexity, is something that can be reasoned about. It's picking the next token, until it hits an EOS. The semantics imposed on those tokens (reasoning ,tool call, etc.) are up to the user('s harness) to decide and act on. The more that's pushed behind the facade, the harder it is achieve sufficient understanding of the model's behavior s.t. one can compose it into larger abstractions. Perhaps the performance (and the adherence to an interface/contract) compensate? But swapping from Opus or 5.5 to this or Fugu seems like a much bigger change than swapping between different 'base' models.
I might be wrong, but strongly suspect that Fable 5 is already something in this shape, considering long time to first token while having normal troughput.
Solutions like these are really cementing the view that LLMs are becoming a commodity
> The phrase "frontier model" is starting to mean two things. One is a checkpoint. The other is a system boundary.
LLM-isms aside, I don't think we want this to be the case? An LLM, for all its complexity, is something that can be reasoned about. It's picking the next token, until it hits an EOS. The semantics imposed on those tokens (reasoning ,tool call, etc.) are up to the user('s harness) to decide and act on. The more that's pushed behind the facade, the harder it is achieve sufficient understanding of the model's behavior s.t. one can compose it into larger abstractions. Perhaps the performance (and the adherence to an interface/contract) compensate? But swapping from Opus or 5.5 to this or Fugu seems like a much bigger change than swapping between different 'base' models.
I might be wrong, but strongly suspect that Fable 5 is already something in this shape, considering long time to first token while having normal troughput.
Can we please stop submitting fully AI-generated text to HN?
at least 50% of the front page would disappear if this were enforced
I'd be perfectly okay with that.
This should help with better utilizing a heterogenous collection of inference hardware.