Researchers posted the preprint "Cut Binary Cross Entropy: Efficient Large-Vocabulary Loss and Gradient Kernels for Sequential Recommendation", which describes CutBCE as "an exact, hardware-accelerated BCE loss and gradient operator implemented in JAX and Pallas for large-vocabulary workloads," according to the abstract on arXiv CS.LG.

The paper targets a practical bottleneck in recommender training. In the abstract, the authors write that industrial sequential recommender systems can operate over item catalogs of "10^5--10^7 items" and that standard binary cross-entropy "materializes a dense [B, N, V] logits tensor in High Bandwidth Memory (HBM), incurring prohibitive O(BNV) memory and fatal Out-Of-Memory (OOM) errors" arXiv CS.LG.

According to the preprint abstract, the proposed method is "an exact" operator rather than an approximation. The authors say CutBCE introduces "an exact fused reformulation evaluating dense background loss and sparse target corrections," plus "a custom Vector-Jacobian Product (VJP) with a dedicated Pallas TPU backward kernel computing logit tiles on-chip in both passes so logits and their gradients never reside in HBM," along with "dynamic VMEM budgeting and sharding-aware collective hoisting for distributed meshes" arXiv CS.LG.

The reported gains are limited to the authors' own tests in the preprint, which is not peer reviewed. On single-chip TPU v5e/v6e mini-benchmarks, the abstract says CutBCE "eliminates OOM errors with up to 91.9% speedup." On 8-chip TPU slice training for multi-label SASRec with 876k items on Yambda-50M, it says CutBCE "reduces peak HBM by 65.7% (>14 GiB saved per chip) and increases training speed by 225.9% with comparable accuracy" arXiv CS.LG.

Those figures come from the preprint authors and, from the abstract alone, do not establish methodology details, variance, or independent replication. The arXiv entry also says CutBCE is open-sourced at "this https URL," but the abstract page itself does not provide the repository destination in the visible text on arXiv.

The paper is listed on arXiv as arXiv:2610.05559 under machine learning, with additional subject tags for artificial intelligence, hardware architecture and information retrieval arXiv CS.LG.