United We TransformCreate teamsGrade your agenda
Atlas/Events/USENIX 2025
action-oriented convening agenda analysis

USENIX 2025

This action-oriented convening in Technology / AI / Startup shows 54 visible agenda rows from usenix.org and scores 50/100: a promising public-evidence design, still below the strong-evidence threshold. The clearest public signals sit in Future-of-Work Fit and Participation Architecture; the main limits are Follow Through and Evidence Maturity. Visible mechanisms include Participant work, Commitments, Feedback, and Network design. Follow-through or tracking is at least visible enough to inspect, though causal... A practical reading: For a reader, this is a useful but still incomplete public example: it reads as an action-oriented convening, with the strongest visible signal in future-of-work fit and participation architecture and the biggest open question around follow through and evidence maturity. The practical test is whether the published agenda connects the room to post-event continuation and evidence. This page is an original public-evidence analysis, not a copy of the source agenda or an endorsement of the event. The score places the visible agenda in the promising but still evidence-limited design band. The strongest visible pillars are Future-of-Work Fit, Participation Architecture, and Learning Transfer; the thinnest visible pillars are Follow Through, Evidence Maturity, and Personalization. Visible mechanisms include Participant work, Commitments, Feedback, Network design, and Learning transfer. The extracted agenda preview includes 80 visible rows. The most common formats are Training, Demo, and Poster Session; the most common inferred purposes are Skill Building, Showcase, and Wellbeing.

Primary source evidence: usenix.org ↗

Eight-pillar fingerprint

Hover any pillar to see what it measures and, where it scored low, what the agenda is missing.

Participation Architecture?73
Participation Architecture - 73/100. Participant work, contribution, interaction, and alternatives to passive broadcast.
Follow Through?16
Follow Through - 16/100. Owners, dates, commitments, progress checks, and accountability after the room.Missing: Add named owners, dates, implementation checkpoints, and a visible post-event continuation path.
Problem Specificity?52
Problem Specificity - 52/100. A clear costly problem, objective, decision, or performance target.
Personalization?49
Personalization - 49/100. Role, path, goal, preparation, or connection tailoring for participants.Missing: Create role-based paths, prepared questions, tailored breakouts, or participant-specific next steps.
Network Design?54
Network Design - 54/100. Structured weak ties, bridge-building, mixers, and relationship persistence.
Learning Transfer?67
Learning Transfer - 67/100. Applied practice, feedback, workplace use, refreshers, and 30-90 day transfer.
Evidence Maturity?35
Evidence Maturity - 35/100. Baseline, comparison, follow-up, isolation, and attribution confidence.Missing: Add baseline measurement, comparison logic, tracking, or post-event impact reporting so effectiveness is not inferred only from format.
Future-of-Work Fit?76
Future-of-Work Fit - 76/100. Value against time, hybrid reality, accessibility, AI, and meeting load.

Fix the gaps

Field-tested exercises matched to this agenda's weakest pillars, from the exercise library.

Agenda Preview

The actual agenda we captured. Every block is classified by format and purpose. Open any block to see how we read it; the colored edge shows whether it is participant work, broadcast, logistics, or a showcase.

Room vs wrapper

56 percent of the 80 classified blocks put participants to work; the rest broadcast, show, or handle logistics. That mix is what drives the participation score.

45
14
6
15
Participant workBroadcastShowcaseLogistics
all eventPoster SessionPoster sessionShowcase+
Format · ShowcasePoster sessionPresenters display work; attendees browse and ask questions. Some interaction, not structured work.
Evidence basisOutcome inferred from formatInferred from format
7:30 am - 9:00 amContinental BreakfastMealPacing+
Format · LogisticsMealA pacing block. Can carry unstructured networking, not scored as participant work.
Evidence basisOutcome inferred from formatInferred from format
9:00 am - 10:00 amHide details ▾PresentationKnowledge transfer+
Format · BroadcastPresentationSpeakers present, the audience receives. Awareness only unless paired with practice or follow-up.
Evidence basisNo participant output visible from this rowRead from source, no work signal
all eventUSENIX ATC '25 and OSDI '25 Joint Keynote AddressKeynoteExpert framing+
Format · BroadcastKeynoteA featured talk from the stage. Builds awareness and energy, produces no participant output on its own.
Evidence basisNo participant output visible from this rowRead from source, no work signal
10:00 am - 10:30 amCoffee and Tea BreakBreakPacing+
Format · LogisticsBreakA pacing or recovery block between sessions.
Evidence basisOutcome inferred from formatInferred from format
10:30 am - 10:45 amOpening Remarks and AwardsOpeningOrientation+
Format · BroadcastOpeningFormat not classified from the source; treated as a broadcast block by default.
Evidence basisNo participant output visible from this rowRead from source, no work signal
10:45 am - 12:25 pmTrack 1PresentationKnowledge transfer+
Format · BroadcastPresentationSpeakers present, the audience receives. Awareness only unless paired with practice or follow-up.
Evidence basisNo participant output visible from this rowRead from source, no work signal
all eventAccelerating ML Training: Parallelism, Tuning, and ModalitiesTrainingSkill building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisParticipant work is implied by the formatInferred from format
all eventFlexPipe: Maximizing Training Efficiency for Transformer-based Models with Variable-Length InputsTrainingSkill building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisParticipant work is implied by the formatInferred from format
all eventOptimus: Accelerating Large-Scale Multi-Modal LLM Training by Bubble ExploitationTrainingSkill building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisParticipant work is implied by the formatInferred from format
12:25 pm - 2:00 pmConference LuncheonMealPacing+
Format · LogisticsMealA pacing block. Can carry unstructured networking, not scored as participant work.
Evidence basisOutcome inferred from formatInferred from format
all eventNetworking: From Cloud to In-Network IntelligenceNetworkingRelationship building+
Format · LogisticsNetworkingUnstructured mixing. Can carry incidental connection, but is not scored as designed network work.
Evidence basisOutcome inferred from formatInferred from format
6:00 pm - 7:30 pmOSDI '25 Poster Session and ReceptionReceptionShowcase+
Format · LogisticsReceptionA social or hospitality block. Pacing and informal connection, not participant work.
Evidence basisOutcome inferred from formatInferred from format
7:30 pm - 8:30 pmUSENIX 50th Anniversary CelebrationPresentationKnowledge transfer+
Format · BroadcastPresentationSpeakers present, the audience receives. Awareness only unless paired with practice or follow-up.
Evidence basisNo participant output visible from this rowRead from source, no work signal
all eventObscura: Concealing Recomputation Overhead in Training of Large Language Models with Bubble-filling Pipeline TransformationTrainingSkill building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisParticipant work is implied by the formatInferred from format
all eventGREYHOUND: Hunting Fail-Slows in Hybrid-Parallel Training at ScaleTrainingSkill building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisParticipant work is implied by the formatInferred from format
all eventPrimus: Unified Training System for Large-Scale Deep Learning Recommendation ModelsTrainingSkill building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisParticipant work is implied by the formatInferred from format
all eventPopFetcher: Towards Accelerated Mixture-of-Experts Training Via Popularity Based Expert-Wise PrefetchTrainingSkill building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisParticipant work is implied by the formatInferred from format
all eventHypeReca: Distributed Heterogeneous In-Memory Embedding Database for Training Recommender ModelsTrainingSkill building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisParticipant work is implied by the formatInferred from format
all eventCrossPipe: Towards Optimal Pipeline Schedules for Cross-Datacenter TrainingTrainingSkill building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisParticipant work is implied by the formatInferred from format
6:00 pm - 7:30 pmUSENIX ATC '25 Poster Session and ReceptionReceptionShowcase+
Format · LogisticsReceptionA social or hospitality block. Pacing and informal connection, not participant work.
Evidence basisOutcome inferred from formatInferred from format
7:30 pm - 8:30 pmTribute to USENIX ATCPresentationKnowledge transfer+
Format · BroadcastPresentationSpeakers present, the audience receives. Awareness only unless paired with practice or follow-up.
Evidence basisNo participant output visible from this rowRead from source, no work signal
all eventAccelerating Model Training on Ascend Chips: An Industrial System for Profiling, Analysis and OptimizationTrainingSkill building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisParticipant work is implied by the formatInferred from format
12:30 pm - 2:00 pmLunch (on your own)MealPacing+
Format · LogisticsMealA pacing block. Can carry unstructured networking, not scored as participant work.
Evidence basisOutcome inferred from formatInferred from format
all eventColocating ML Inference and Training with Fast GPU Memory HandoverTrainingSkill building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisParticipant work is implied by the formatInferred from format
5:30 pm - 5:35 pmClosing RemarksClosingOrientation+
Format · BroadcastClosingFormat not classified from the source; treated as a broadcast block by default.
Evidence basisNo participant output visible from this rowRead from source, no work signal
all eventPoster SessionPoster SessionShowcase+
Format · ShowcasePoster SessionPresenters display work; attendees browse and ask questions. Some interaction, not structured work.
Evidence basisMediumRead from source
7:30 am - 9:00 amContinental BreakfastMealWellbeing+
Format · LogisticsMealA pacing block. Can carry unstructured networking, not scored as participant work.
Evidence basisMediumRead from source
9:00 am - 10:00 amHide details ▾UnknownUnknown+
Format · BroadcastUnknownFormat not classified from the source; treated as a broadcast block by default.
Evidence basisMediumRead from source
all eventUSENIX ATC '25 and OSDI '25 Joint Keynote AddressKeynoteThought Leadership+
Format · BroadcastKeynoteA featured talk from the stage. Builds awareness and energy, produces no participant output on its own.
Evidence basisMediumRead from source
10:00 am - 10:30 amCoffee and Tea BreakBreakWellbeing+
Format · LogisticsBreakA pacing or recovery block between sessions.
Evidence basisMediumRead from source
10:30 am - 10:45 amOpening Remarks and AwardsOpening RemarksOrientation+
Format · BroadcastOpening RemarksFraming or welcome from the stage. Orients the room, not participatory.
Evidence basisMediumRead from source
10:45 am - 12:25 pmTrack 1UnknownUnknown+
Format · BroadcastUnknownFormat not classified from the source; treated as a broadcast block by default.
Evidence basisMediumRead from source
all eventIn this paper, we propose DEEPSERVE, a scalable and serverless AI platform designed to efficiently serve large language models (LLMs) at scale in cloud environments. DEEPSERVE addresses key challenges such as resource allocation, serving efficiency, and cold start latencies through four main design components. First, DEEPSERVE uses a simple serverless abstraction called the request-job-task model, which helps manage diverse AI workloads across post-training and model-serving tasks.TrainingSkill Building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisMediumRead from source
all eventAccelerating ML Training: Parallelism, Tuning, and ModalitiesTrainingSkill Building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisMediumRead from source
all eventFlexPipe: Maximizing Training Efficiency for Transformer-based Models with Variable-Length InputsTrainingSkill Building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisMediumRead from source
all eventOptimus: Accelerating Large-Scale Multi-Modal LLM Training by Bubble ExploitationTrainingSkill Building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisMediumRead from source
all eventThis paper proposes Optimus, a distributed MLLM training system that reduces end-to-end MLLM training time. Optimus is based on our principled analysis that scheduling the encoder computation within the LLM bubbles can reduce bubbles in MLLM training.TrainingSkill Building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisMediumRead from source
all eventTo enable scheduling encoder computation for all GPUs, Optimus searches for separate parallel plans for the encoder and LLM, and adopts a bubble scheduling algorithm to exploit LLM bubbles without breaking the original data dependencies in the MLLM model architecture. We further decompose the encoder layer computation into a series of kernels and analyze the common bubble pattern of 3D parallelism to carefully optimize the sub-millisecond bubble scheduling, minimizing the overall training time. Our experiments in a production cluster show that Optimus accelerates MLLM training by 20.5%-21.3% with ViT-22B and GPT-175B model over 3072 GPUs compared to baselines.TrainingSkill Building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisMediumRead from source
12:25 pm - 2:00 pmConference LuncheonMealWellbeing+
Format · LogisticsMealA pacing block. Can carry unstructured networking, not scored as participant work.
Evidence basisMediumRead from source
all eventNetworking: From Cloud to In-Network IntelligenceNetworkingRelationship Building+
Format · LogisticsNetworkingUnstructured mixing. Can carry incidental connection, but is not scored as designed network work.
Evidence basisMediumRead from source
all eventTo address this issue, we propose MARC, a motion-aware rate control framework that aligns bitrate decisions with user QoE preferences in real-time. MARC sets dynamic QoE objectives based on real-world user engagement behavior, captures the different latency and quality requirements for motion and non-motion frames, and employs stochastic optimization to maximize QoE. Extensive deployment of over 1 million user sessions demonstrates that MARC reduces session freeze rates by 71% and increases user interaction time by 20%, significantly improving user engagement for e-commerce cloud rendering.DemoShowcase+
Format · Participant workDemoA hands-on or applied walkthrough that invites attendee questions and direct engagement.
Evidence basisMediumRead from source
all eventThis paper presents GPreempt, a preemption mechanism that breaks the trade-off. GPreempt implements a timeslice-based yield mechanism to enable context-switch preemption on GPUs. To mitigate the overhead associated with context-switching, GPreempt employs a hint-based pre-preemption technique to overlap the preemption process with the essential data-preparation phase. Our evaluation demonstrates that GPreempt achieves within 40 μs low-latency preemption comparable to executing only latency-critical tasks while remaining applicable to non-idempotent workloads, where reset-based mechanisms prove inadequate.DemoShowcase+
Format · Participant workDemoA hands-on or applied walkthrough that invites attendee questions and direct engagement.
Evidence basisMediumRead from source
all eventWe propose μEFI, the first isolation framework for UEFI firmware that can transparently run UEFI modules in sandboxes. Drawing inspiration from microkernel design, we deprivilege UEFI modules to user mode and isolate them in different address spaces (sandboxes). To enable the transparent execution of UEFI modules, we propose trampoline injection and protocol analysis. To further strengthen UEFI security, we incorporate a seccomp-like mechanism to restrict module capabilities and perform automated input validation to detect and prevent invalid inputs. Evaluation results demonstrate that our system can run complex UEFI modules without modifications, which incurs a small overhead of 1.91% for UEFI boot phase.DemoShowcase+
Format · Participant workDemoA hands-on or applied walkthrough that invites attendee questions and direct engagement.
Evidence basisMediumRead from source
all eventTo increase platform memory efficiency, hyperscalers like Google and Meta transparently demote "cold" application data to cheaper cost-per-byte memory tiers like compressed memory and NVMe SSDs. These systems rely on standard kernel paging policies and mechanisms to maximize the achievable memory savings without hurting application performance. Although the literature promises better policies, implementing and deploying them within the Linux kernel is challenging. Delegating policies and mechanisms to user space, through userfaultfd or library-based approaches, incurs overheads and may require modifying application code.DemoShowcase+
Format · Participant workDemoA hands-on or applied walkthrough that invites attendee questions and direct engagement.
Evidence basisMediumRead from source
all eventSafe kernel extensions have gained significant traction, evolving from simple packet filters to large, complex programs that customize storage, networking, and scheduling. Existing kernel extension mechanisms like eBPF rely on in-kernel verifiers to ensure safety of kernel extensions by static verification using symbolic execution. We identify significant usability issues - safe extensions being rejected by the verifier - due to the language-verifier gap, a mismatch between developers’ expectation of program safety provided by a contract with the programming language, and the verifier’s expectation.NetworkingRelationship Building+
Format · LogisticsNetworkingUnstructured mixing. Can carry incidental connection, but is not scored as designed network work.
Evidence basisMediumRead from source
all eventEnd-to-end deep neural networks' (DNNs) training performance depends not only on the time spent in training the model weights but also on the time spent in loading and preprocessing the training data. Recent advances in GPU hardware have made training substantially faster. As a result, the bottleneck has shifted to the CPU-based input pipeline. This pipeline must fetch and transform each sample through multiple stages before it can be consumed by the GPU.TrainingSkill Building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisMediumRead from source
6:00 pm - 7:30 pmOSDI '25 Poster Session and ReceptionPoster SessionShowcase+
Format · ShowcasePoster SessionPresenters display work; attendees browse and ask questions. Some interaction, not structured work.
Evidence basisMediumRead from source
all eventWould you like to share a provocative opinion, interesting preliminary work, or a cool idea that will spark discussion at this year's OSDI? The poster session is the perfect venue to introduce such new or ongoing work. Poster presenters will have the opportunity to discuss their work, get exposure, and receive feedback from other attendees during the in-person evening reception. View the list of accepted posters.Poster SessionShowcase+
Format · ShowcasePoster SessionPresenters display work; attendees browse and ask questions. Some interaction, not structured work.
Evidence basisMediumRead from source
7:30 pm - 8:30 pmUSENIX 50th Anniversary CelebrationUnknownUnknown+
Format · BroadcastUnknownFormat not classified from the source; treated as a broadcast block by default.
Evidence basisMediumRead from source
all eventWe propose a mechanism called workload weaving , which offloads attention operators of hot models to running cold models, achieving high GPU memory utilization with low communication cost. To mitigate the blocking caused by running cold models, we propose WEAVER with two key techniques: (i) GPU-driven dynamic control flow, which delegates the control logic of offloading to GPUs, letting the offloaded operators bypass pending kernels in the GPU hardware queue; (ii) operator splitting, which carefully divides the large kernels of cold models into smaller ones to mitigate the head-of-line blocking. Our evaluation using real-world LLM trace demonstrates that WEAVER improves the throughput of hot models by up to 77% while maintaining the same or lower TPOT. For the cold model, WEAVER incurs a modest overhead (3-5ms).DemoShowcase+
Format · Participant workDemoA hands-on or applied walkthrough that invites attendee questions and direct engagement.
Evidence basisMediumRead from source
all eventObscura: Concealing Recomputation Overhead in Training of Large Language Models with Bubble-filling Pipeline TransformationTrainingSkill Building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisMediumRead from source
all eventPipeline parallelism has become a widely adopted strategy for training large language models (LLMs) by distributing computational workloads across multiple nodes. However, it faces a significant challenge in the form of memory bottlenecks at early stages. While recomputation can mitigate this issue, it incurs additional computational overhead.TrainingSkill Building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisMediumRead from source
all eventTo address this limitation, we propose Obscura, a computationally efficient pipeline training system designed to optimize recomputation overhead under the given memory constraints. Leveraging the observation that bubbles following backward passes can conceal recomputation overhead in pipeline parallelism, Obscura introduces a novel pipeline transformation to enhance overhead concealment. Furthermore, we integrate swapping techniques into the pipeline and model the execution time as an optimization problem to identify an optimal recomputation strategy. A partition adjustment algorithm is also implemented to balance computation across stages under the transformation. Evaluations on Llama-2 and GPT-3 models of various sizes demonstrate that Obscura achieves throughput improvements of up to 1.33× compared to widely used recomputation baselines.TrainingSkill Building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisMediumRead from source
all eventZhaoyang Wan, Rongxin Han, Haifeng Sun, Qi Qi, Zirui Zhuang, and Bo He, State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications; Liang Zhang, Huawei Technologies Co., Ltd; Jianxin Liao and Jingyu Wang, State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and TelecommunicationsNetworkingRelationship Building+
Format · LogisticsNetworkingUnstructured mixing. Can carry incidental connection, but is not scored as designed network work.
Evidence basisMediumRead from source
all eventGREYHOUND: Hunting Fail-Slows in Hybrid-Parallel Training at ScaleTrainingSkill Building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisMediumRead from source
all eventByte-addressable non-volatile memory (NVM) proposes a new opportunity to enable file systems with better performance and durability by adding a new persistent caching layer. However, crash consistency of caching layers and compatibility with diverse reliability features of file systems remains unknown. This paper conducts a crash consistency case study of Open CAS, a popular block-level caching system. Through careful and thorough crash consistency experiments, we show that Open CAS cannot always maintain crash consistency in the persistent caching layer. We also demonstrate some reliability features of file systems are not compatible with Open CAS. Our analysis reveals the importance of a systematic crash consistency test to caching systems and reliability feature co-design with file systems in the construction of a reliable end-to-end file system.DemoShowcase+
Format · Participant workDemoA hands-on or applied walkthrough that invites attendee questions and direct engagement.
Evidence basisMediumRead from source
all eventPrimus: Unified Training System for Large-Scale Deep Learning Recommendation ModelsTrainingSkill Building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisMediumRead from source
all eventPrimus has demonstrated its efficiency and effectiveness in handling large-scale, enterprise-grade DLRM training over five years of deployment at ByteDance. Evaluations show Primus’s optimizations of resources, data, and paradigms. Firstly, dynamic scaling reduces training cost by 17.1% at the cluster level and increases CPU utilization from 50% to 80% per job. Secondly, data orchestration accelerates task generation by 23× and achieves higher training throughput. Lastly, after applying the hybrid training paradigm with 4 different DLRMs, advertising revenue increases by 0.4%-2.4%.TrainingSkill Building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisMediumRead from source
all eventPopFetcher: Towards Accelerated Mixture-of-Experts Training Via Popularity Based Expert-Wise PrefetchTrainingSkill Building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisMediumRead from source
all eventHypeReca: Distributed Heterogeneous In-Memory Embedding Database for Training Recommender ModelsTrainingSkill Building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisMediumRead from source
all eventMaking high-quality recommendations is important in online applications. To improve user satisfaction and effectiveness of advertising, deep learning-based recommender models (DLRM) are widely studied and deployed. Training these models on massive data demands increasing computation power, commonly provided by a cluster of numerous GPUs. Meanwhile, the embedding tables of the models are huge, posing challenges on the memory. Existing systems exploit host memory and hashing techniques to accommodate them. However, the simple offloading design is hard to scale up to multiple nodes. The sparse access to the distributed embedding tables introduces high data management and all-to-all communication overhead.TrainingSkill Building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisMediumRead from source
all eventWe find that a distributed in-memory key-value database is the best abstraction to serve and maintain embedding vectors in DLRM training. To achieve high scalability, our system, HypeReca, utilizes both GPU and CPU memory. We improve the throughput of data management according to the batching pattern of DNN training, using a pipeline over decentralized indexing tables and a contentionavoiding schedule for data exchange. A two-fold parallel strategy is used to guarantee consistency of all embedding vectors. The communication overhead is reduced by replicating a few frequently accessed embedding vectors, exploiting the sparse pattern with a performance model. In our evaluation on 32 GPUs over real-world datasets, HypeReca achieves 2.16 to 16.8× end-to-end speedup over HugeCTR, TorchRec and TFDE. The source code is available at <https://github.com/thu-pacman/hypereca/>.TrainingSkill Building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisMediumRead from source
all eventCrossPipe: Towards Optimal Pipeline Schedules for Cross-Datacenter TrainingTrainingSkill Building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisMediumRead from source
6:00 pm - 7:30 pmUSENIX ATC '25 Poster Session and ReceptionPoster SessionShowcase+
Format · ShowcasePoster SessionPresenters display work; attendees browse and ask questions. Some interaction, not structured work.
Evidence basisMediumRead from source
all eventThe USENIX ATC '25 poster session and reception will feature posters by authors presenting their work at the conference. View the list of accepted posters.Poster SessionShowcase+
Format · ShowcasePoster SessionPresenters display work; attendees browse and ask questions. Some interaction, not structured work.
Evidence basisMediumRead from source
7:30 pm - 8:30 pmTribute to USENIX ATCUnknownUnknown+
Format · BroadcastUnknownFormat not classified from the source; treated as a broadcast block by default.
Evidence basisMediumRead from source
all eventWe demonstrate effectiveness of HEC on PolyBenchC benchmarks, successfully verifying loop unrolling, tiling, and fusion transformations. HEC processes over 100,000 lines of MLIR code in 40 minutes with predictable runtime scaling. Importantly, HEC identified two critical compilation errors in mlir-opt: loop boundary check errors causing unintended executions during unrolling, and memory read-after-write violations in loop fusion that alter program semantics. These findings demonstrate HEC practical value in detecting real-world compiler bugs and highlight the importance of formal verification in optimization pipelines.DemoShowcase+
Format · Participant workDemoA hands-on or applied walkthrough that invites attendee questions and direct engagement.
Evidence basisMediumRead from source
all eventThis paper proposes an efficient PIM-capable ANNS system named PIMANN. We observe that each PIM core has an additional, undocumented, and little-known control interface (originally used for control commands like launching PIM kernels), which could be retrofitted for fine-grained arbitration of DDR bus access. Thus, PIMANN can break the traditional batching scheduling paradigm and adopt a fine-grained, per-PIM-core scheduling paradigm. With this key idea, PIMANN introduces 1) persistent PIM kernel technique to eliminate the idle state between two batches, and 2) per-PU query dispatching technique that dispatches queries based on the real-time load of PIM cores. Experiments show that PIMANN can boost throughput by 2.4-10.4× compared to existing ANNS systems on CPU or GPU. The implementation of PIMANN is available at <https://github.com/cds-ruc/PIM-ANNS>.BreakWellbeing+
Format · LogisticsBreakA pacing or recovery block between sessions.
Evidence basisMediumRead from source
all eventWe prototype SwCC using the Xilinx U280 FPGA. Experimental results demonstrate that SwCC achieves performance comparable to current commercial RDMA NICs (Mellanox ConnextX-5). Both SwCC and ConnectX-5 reach 3.1 µs control loop RTT and need 512B packet size to reach line-rate traffic (100 Gbps). In terms of flexibility, SwCC allows to use the C language to implement nearly all kinds of existing CCAs, e.g., rate-based CCAs, window-based CCAs, and credit-based CCAs. The potential ASIC design of SwCC can easily scale to higher network bandwidth.DemoShowcase+
Format · Participant workDemoA hands-on or applied walkthrough that invites attendee questions and direct engagement.
Evidence basisMediumRead from source
all eventWe therefore propose SpaceExit, an integrated system for efficient adaptive computing on satellites. SpaceExit introduces three key components: (1) a geospatial-contextual adaptive detector that leverages both visual semantics and geospatial context to adjust processing complexity for each image, (2) a complexity-driven adaptive task scheduler that partitions images into tiles and allocates inference tasks across onboard devices based on content complexity and device capabilities, and (3) a satellite resource adaptive controller that ensures safe and efficient execution under changing conditions. Evaluations of diverse satellite settings and hardware platforms demonstrate that SpaceExit increases the performance by 5.2%-37.6% compared with the SoTA design.DemoShowcase+
Format · Participant workDemoA hands-on or applied walkthrough that invites attendee questions and direct engagement.
Evidence basisMediumRead from source
all eventAccelerating Model Training on Ascend Chips: An Industrial System for Profiling, Analysis and OptimizationTrainingSkill Building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisMediumRead from source
12:30 pm - 2:00 pmLunch (on your own)MealWellbeing+
Format · LogisticsMealA pacing block. Can carry unstructured networking, not scored as participant work.
Evidence basisMediumRead from source
all eventThis paper explores a class of important atomicity properties between the container-like arrays and their logical size variables, referred to as the counting correlation, which are common in PM programs but exceed the capability of existing approaches. We propose invariants to capture the necessary behaviors of counting-correlated variables, utilize symbolic range analysis to extract PM program behaviors, and encode them into SMT constraints. These constraints are checked against the invariants to infer likely PM program properties. We demonstrate the utility of the inferred properties by leveraging them for PM bug detection, which discovers 14 atomicity bugs (including 11 new bugs) in real-world PM programs.DemoShowcase+
Format · Participant workDemoA hands-on or applied walkthrough that invites attendee questions and direct engagement.
Evidence basisMediumRead from source
all eventIn this paper, we present 2DFS, a novel two-dimensional filesystem that enables independent updates, caching, and distribution of ML model components. We design and develop a complete ecosystem, including a builder, registry, and cache hierarchy, to streamline the build and deployment processes of ML models leveraging 2DFS. Our comprehensive evaluation of 14 real-world ML models demonstrates that 2DFS achieves up to 56x faster build times, 25x better caching efficiency, while providing on-demand image partitioning with negligible overhead. 2DFS is fully OCI-compliant and integrates seamlessly with existing infrastructures and container workflows.DemoShowcase+
Format · Participant workDemoA hands-on or applied walkthrough that invites attendee questions and direct engagement.
Evidence basisMediumRead from source
all eventUniversal Checkpointing: A Flexible and Efficient Distributed Checkpointing System for Large-Scale DNN Training with Reconfigurable ParallelismTrainingSkill Building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisMediumRead from source
all eventDeep neural network (DNN) training continues to scale rapidly in terms of model size, data volume, and sequence length, to the point where multiple machines are required to fit large models for training. Different distributed and parallel training strategies have been developed to support large-scale DNN training by partitioning the training state across GPUs. However, existing DNN training systems provide very limited support for reconfiguring parallelism strategies in the middle of the training via checkpointing. This limitation arises because distributed checkpoints are tightly coupled to specific model parallelism and hardware configurations, preventing large-scale training jobs from efficiently adapting to hardware failures or resource elasticity.TrainingSkill Building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisMediumRead from source
all eventThis paper presents Universal Checkpointing (UCP), a novel checkpointing system that enables flexible and efficient DNN training with reconfigurable parallelism. UCP overcomes challenges in existing systems by decoupling checkpoint structure from parallel training strategies and hardware configurations. In addition, we present a pattern-based reconfiguration pipeline that enables automatic, flexible, and efficient mapping of checkpoint state to various parallelism strategies. Evaluation on a range of DNN models, including state-of-the-art dense and sparse LLMs, shows that UCP enables reconfiguration for a broader set of widely used parallelism strategies than existing solutions while adding negligible reconfiguration cost. UCP has been successfully employed in real LLM training workloads, greatly enhancing their flexibility and resilience to dynamic hardware environments.TrainingSkill Building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisMediumRead from source
all eventColocating ML Inference and Training with Fast GPU Memory HandoverTrainingSkill Building+
Format · Participant workTrainingGuided skill building where participants practice. Counts as participant work and learning transfer.
Evidence basisMediumRead from source
5:30 pm - 5:35 pmClosing RemarksClosing RemarksOrientation+
Format · BroadcastClosing RemarksClosing framing from the stage. Wraps the event, not participatory.
Evidence basisMediumRead from source

The Full Reading

Why It Ranks This Way +

Calibrated from GES design 54/100 and verified 53/100 with no fourth-loop cap.

Reader Takeaway. For a reader, this is a useful but still incomplete public example: it reads as an action-oriented convening, with the strongest visible signal in future-of-work fit and participation architecture and the biggest open question around follow through and evidence maturity. The practical test is whether the published agenda connects the room to post-event continuation and evidence.

Strongest signals: Future-of-Work Fit, Participation Architecture, and Learning Transfer. Weakest signals: Follow Through, Evidence Maturity, and Personalization.

How This Agenda Could Improve +
  • Add named owners, dates, implementation checkpoints, and a visible post-event continuation path.
  • Add baseline measurement, comparison logic, tracking, or post-event impact reporting so effectiveness is not inferred only from format.
  • Create role-based paths, prepared questions, tailored breakouts, or participant-specific next steps.

Fastest next move: Add named owners, dated next steps, and a visible continuation path before treating the event as outcome-ready.

Role-Specific Reading +

Event owner lens

Use this record to benchmark whether a comparable event makes the work after the room visible. The score is 50/100, so the next move is to benchmark the weakest pillars before repeating the format.

Sponsor lens

Look beyond exposure. Strong sponsor value would show qualified interaction, problem work, buyer learning, customer evidence, or follow-up. The practical sponsor move is to look for structured introductions, buyer-seller fit, and relationship persistence.

Designer lens

The agenda is useful as a pattern sample from usenix.org. Redesign attention should go first to the lowest-scoring pillars; in practice, turn the thinnest agenda blocks into participant work.

Executive lens

Treat the visible agenda as an operating plan. The executive move is to require owners, dates, and evidence before treating the event as strategic. If owners, proof, and follow-through are not visible, the public record does not yet prove strategic movement.

Aggregator lens

Treat the source URL as evidence, not decoration. The data-product move is to label the source boundary clearly before ranking the record before ranking or syndicating the record.

What GES Means Here +

The Gathering Effectiveness Score is a strict 0-100 public-evidence reading of the agenda across eight pillars. It rewards visible participant work, follow-through, transfer, network design, and proof mechanisms more than polish, speaker fame, attendance, or satisfaction.

Visible mechanisms: Participant work, Commitments, Feedback, Network design, Learning transfer, Personalization.

Evidence boundary: Scores reflect visible agenda/source evidence and should not be read as proof of causal event impact.

Limitations, Score Caps, and Review Flags +

Limitations

  • No baseline measurement is visible.

Score caps

  • No fourth-loop score cap applied.

Review flags

  • Satisfaction/NPS signal found, but it is excluded from effectiveness scoring.
  • No tracking, validation, feedback, or impact measurement found in the visible source text.
  • High cleanup rate: many extracted rows were hidden or merged as fragments.
  • Original category was unknown; publication category is inferred.
Is this proof the event worked? +

No. This is a strict public-evidence reading of the agenda. Proof would require baseline, comparison, follow-up, attribution, and impact evidence beyond the listing.

What should a reader inspect first? +

Start with the source URL, then compare the eight pillar scores against the agenda rows. The biggest opportunities usually sit in follow-through, evidence maturity, and participant work.

Why publish weak records? +

Weak records are part of the map. They show where public agendas still describe sessions and speakers more often than outcomes, commitments, transfer, or proof.

How should I use the rows? +

Read the agenda rows as the visible design trace: formats, purposes, and evidence labels show what the public source made inspectable, not everything that happened in the room. This is a source-grounded interpretation of the public agenda record, not a copy of the source, and not an endorsement of the event.

Where To Go Next

Compare this agenda against other Technology / AI / Startup events scored on the same eight pillars.