Discussion about this post

User's avatar
Mark's avatar

Given OneChronos' strength in combinatorics, I initially expected a multi-asset hardware portfolio matching - e.g., a portfolio of 2x H100 with 80GB memory, 4x B200 with 180GB, 2x GB200 with 192GB within a NVL72 server rack. But that may lead to too many variations when considering not just the server, interconnect, and memory setup but also latency, TPS, TTFT, etc.

Using workload makes sense, given that it enables a higher level of abstraction that makes standardization easier. I'm now curious what that looks like in the context of TPS and TTFT, or whether that is somehow abstracted away as well, with the workload being some type of a deliverable like finishing a specific task.

Great piece - lots of food for thought.

No posts

Ready for more?