IOTA SDK Trains 16B Model Across 139 GPUs as Macrocosmos Reports Early Results
Macrocosmos has published its first technical results from the IOTA SDK, just one week after announcing the distributed AI software at Exploit Summit. The company reports that its team has already used the SDK across pretraining, reinforcement learning and inference workloads, providing an early look at the performance of its distributed training stack.
One of the largest tests involved training a 16-billion-parameter model across 139 consumer GPUs comprising NVIDIA RTX 4090 and RTX 5090 hardware. Macrocosmos reported a 1.8x wall-time speedup compared with its legacy codebase, indicating that the new SDK was able to reduce the time required for the workload under the tested configuration.
IOTA SDK Delivers Distributed Training Gains
Macrocosmos also reported a 30% reduction in Butterfly all-reduce times. The company attributed the improvement to new abstractions within the IOTA SDK, which are designed to simplify and improve communication across distributed GPU systems.
The training results extend beyond the 16B model test. On the same 2B configuration, Macrocosmos said the IOTA SDK achieved 2.3 times the model FLOPs utilization, or MFU, of its previous stack. MFU is a measure of how efficiently available GPU computing capacity is being used during model training.
The company also tested a 30B mixture-of-experts model through ablation experiments on B300 GPUs using autoresearch. Mixture-of-experts architectures divide workloads among specialized groups of parameters, making distributed computing performance an important factor when scaling these systems.
Macrocosmos’ tests were not limited to pretraining. The team reported running GPT-OSS-120B across A6000 GPUs at 878 tokens per second for 32 users, highlighting the SDK’s application to inference workloads involving multiple concurrent users.
The company also served a 320B MoE model across nine GPUs using speculative decoding. That approach can accelerate inference by allowing a smaller model or prediction process to propose tokens before a larger model verifies them, potentially increasing generation speed.
These tests provide an early indication that Macrocosmos is attempting to use the IOTA SDK across multiple stages of the AI model lifecycle rather than focusing exclusively on distributed training. Pretraining, reinforcement learning and inference have different technical requirements, making performance across all three areas significant for the project’s development.
IOTA SDK Expands Into Distributed AI Workloads
Macrocosmos also reported its first results from fully distributed reinforcement learning using LoRA and GRPO on a 7B model. The learner was distributed across three A6000 GPUs, while three additional A6000 GPUs operated as rollout workers.
In the reported MATH-500 benchmark, the setup improved performance from 63.6% to 68.4%. The figures represent Macrocosmos’ stated results from its early testing and provide an indication of how the distributed reinforcement learning setup performed under the company’s configuration.
LoRA, or Low-Rank Adaptation, is commonly used to fine-tune large models more efficiently by updating a smaller set of parameters. GRPO, meanwhile, is a reinforcement-learning approach that can be used to improve model behavior against defined rewards or evaluation criteria.
Taken together, the results cover three major areas of AI infrastructure: model pretraining, reinforcement learning and inference. Macrocosmos says all of these experiments were conducted using the IOTA SDK, making the results an early technical demonstration of the system’s intended scope.
The company has also acknowledged that the SDK remains a work in progress. Macrocosmos said there is still significant work to fix and improve, meaning the reported benchmarks should be viewed as early results rather than final performance figures for a mature platform.
Still, the range of tests gives the IOTA SDK a substantial technical starting point only a week after its public announcement. Training a 16B model across 139 consumer GPUs, alongside large-scale MoE, inference and distributed RL experiments, shows the team is already testing the software under demanding workloads.
For the IOTA ecosystem, the results add another layer to the project’s growing focus on distributed computing and AI infrastructure. Whether the early performance gains translate into broader adoption will depend on further testing, reproducibility and continued improvements to the SDK as Macrocosmos moves beyond its initial technical readout.














