New📄: Skill learning helps agents adapt to new domains. What about making agents more cost-efficient as well?
We introduce SpeedRunner: the first skill learning paper to make cost a primary optimization target 🧵
Introducing AgentOdyssey — an open-ended, long-horizon text game generation engine for 𝐭𝐞𝐬𝐭-𝐭𝐢𝐦𝐞 𝐜𝐨𝐧𝐭𝐢𝐧𝐮𝐚𝐥 𝐥𝐞𝐚𝐫𝐧𝐢𝐧𝐠 𝐚𝐠𝐞𝐧𝐭𝐬.
Real-world agents cannot have a boundary between training and testing: they must learn continuously from interaction with
Tools break in the real world all the time, but not much attention has been given to how well LLMs deal with tool failures.
We introduce HOHW, a tool-use benchmark where problems remain solvable even when tools break adversarially.
📢 New Preprint:
🤔 Humans often reason using elaborate analogies, but can LLMs do the same? 🧵1/5
We introduce AnaloBench, a benchmark of semantically rich analogies that challenges SOTA LLMs 🤖
📄 arxiv.org/abs/2402.12370