免费领取大会全套演讲PPT    

点击领取

我要参会

Yaowei Zheng

Author of PenguinHarness and LlamaFactory

Yaowei Zheng is the Co-founder and CTO of PrismShadow and a Ph.D. candidate at the School of Computer Science and Engineering, Beihang University. He is also the creator of LlamaFactory, a large language model fine-tuning framework that has received over 70,000 stars on GitHub and is widely used by developers worldwide for efficient LLM fine-tuning and deployment. His research成果 have been adopted in practical applications by companies including Alibaba Cloud and NVIDIA. He has published more than 10 papers at leading conferences and journals, including ACL, CVPR, and AAAI, and has served as a reviewer for conferences such as AAAI and EMNLP. His honors include the Beihang Role Model Award and the Ascend Ecosystem Open Source Outstanding Contribution Award. He has also been invited to deliver keynote speeches at major industry events, including the AI Computing Conference and the Alibaba Cloud Apsara Conference. Currently, he leads the development of PenguinHarness, an open-source self-evolving engine for AI agents, aiming to advance agents from automated construction to automated optimization, enabling them to continuously improve through use in vertical domains.

Topic

Building a Low-Cost, High-Return Self-Evolving Agent Loop: The Technical Evolution from LlamaFactory to PenguinHarness

Self-evolution is a critical capability that enables AI agents to move beyond being merely usable toward becoming increasingly capable through continued use. However, how to achieve stable evolution and how to measure it scientifically remain major challenges for the industry. In this talk, we will share our team's end-to-end exploration and practical experience in self-evolving agents. Starting with LlamaFactory, an open-source framework for efficient LLM fine-tuning, we spent six months developing PenguinHarness, an open-source self-evolution engine for AI agents. We will also introduce GDPevo, the first benchmark designed to evaluate an agent's evolutionary capabilities on economically valuable tasks. The talk will provide an in-depth look at how to build a self-evolution loop, covering data generation, multi-agent evaluation, and automated iteration with Harness. We will compare the real-world performance of few-shot evolution and reflective evolution, demonstrating an average accuracy improvement of 18.8% while reducing Token costs by 7.5%. We will also showcase a practical case in which a complete AI application can be generated with a single sentence at a cost of just RMB 0.2, with real-world deployment costs as low as 1/70 of those of Claude Code + Opus. Finally, we will explore the future of fully automated evaluation and continuous agent self-improvement. Outline Background and Challenges: Why do Agents Need Self-Evolution? Limitations of existing frameworks such as LangGraph and LangChain. PenguinHarness Architecture An Agent-oriented SDK, a minimal toolset and system prompt design, and unified access to more than 1,000 models. Building the Self-Evolution Loop A complete Skills framework for automated iteration, from data generation to evaluation. GDPevo Benchmark How to scientifically measure evolutionary capability, covering 120 tasks across 12 business scenarios, with rule-based crossover mechanisms designed to prevent data leakage. Real-World Results and Case Studies A +18.8% improvement in accuracy, a 7.5% reduction in Token consumption, and a real-world cost comparison demonstrating costs as low as 1/70, along with a live demonstration. Looking Ahead Fully automated evaluation, agents that become increasingly capable through use in vertical domains, and collaborative open-source development. Audience Takeaways Understand the Principles of Self-Evolution Learn how to build a complete self-evolution loop for AI agents—including data generation, evaluation, and Harness-based iteration—and apply it directly to your own business scenarios. Gain a Practical Evaluation Methodology Take away GDPevo as a practical framework for measuring evolutionary capability and learn how to scientifically evaluate whether an Agent is genuinely improving. Discover a Low-Cost Deployment Path Learn about unified access to more than 1,000 models and local deployment options, reducing costs to as low as 1/70 of Claude Code + Opus while also supporting enterprise data security requirements. Get Open-Source Tools Ready to Use Access the open-source codebase of PenguinHarness and its Skills packages to quickly build your own self-evolving Agent loop.

© boolan.com 博览 版权所有

沪ICP备15014563号-6

沪公网安备31011502003949号