HF Daily Papers 2026-07-27
范围:优先 RLHF / preference optimization / alignment 在 Diffusion 或生成式 Diffusion 中的应用;其次为 RLHF 在 VLM / LLM 中的工作。排除 VLA、机器人、具身智能、robot manipulation、coding / software-engineering agents、long-horizon agents 与 agentic-RL。
筛选说明
- 本列表只基于 Hugging Face Daily Papers 可见元数据与摘要完成初筛和排序;不应将作者摘要中的主张视为已被独立验证的事实。
- 中文 AI Summary 是摘要级判断,不等同于全文结论。
- 相关论文数:0。
- 全文精读候选数:0。
排序后的相关论文
当天没有论文满足当前范围。
筛选备注
No papers in the provided list meet the scoped focus criteria (RLHF/preference optimization/post-training applied to Diffusion models, or RLHF/preference optimization/post-training for VLMs/LLMs). The candidate papers focus on data-preparation benchmarks, agentic skill frameworks, PyTorch training frameworks, context management architecture, training control planes, native multimodal pre-training from scratch (without RLHF/post-training focus), one-step generative scattering, cross-encoder reranking, video anomaly detection, and latent agent control interfaces.