OffNet Newsroom

Daily topic roundup

AI / ML

Thursday, July 09, 2026 · 2 stories, curated & summarized — click any story for the source.

Third-party benchmarks reveal that SambaNova's heterogeneous platform, which pairs Nvidia H200 GPUs with SN50 RDUs, achieves 763 tokens per second running the MiniMax M2.7 model. This performance metric suggests the architecture can effectively leverage existing Nvidia hardware while adding custom silicon to enhance inference speed. The results highlight a potential pathway for extending the useful life of current GPU fleets through specialized hybrid compute designs.

  • SambaNova's hybrid approach combines off-the-shelf Nvidia H200s with custom SN50 RDUs.
  • MiniMax M2.7 inference hits 763 tok/s in third-party heterogeneous testing.
  • Architecture demonstrates viable path to extend value of aging GPU inventory.
  • Intel-backed startup positions itself as alternative to pure Nvidia or custom silicon stacks.
InfoQ generaldevops ↺ since 07-07

HubSpot Scales Semantic Search to 20B Vectors for Agents and RAG

HubSpot transformed a semantic search proof of concept into an internal service managing over 20 billion vectors for more than 38 teams. The platform now underpins critical workflows including AI agents, retrieval-augmented generation, and contact deduplication. Rising agent usage has shifted the operational priority toward optimizing retrieval quality and minimizing latency.

  • HubSpot's semantic search handles 20B+ vectors across 38+ internal teams.
  • System supports AI agents, RAG pipelines, and automated contact deduplication.
  • Increasing agent traffic makes retrieval accuracy and low latency critical priorities.
  • Architecture evolved from small POC to large-scale enterprise internal service.