Getting an AI agent to actually work is much more difficult than many companies anticipated, but there is help along the way. New startups are finding better ways to test and train agents before deploying them, especially when it comes to the complexities of modern enterprises.
Arga Labs is one such company, announcing a $10 million seed round on Wednesday. The round was led by General Catalyst with participation from Box Group, Emergence, Gradient, and SV Angel.
Arga Labs builds training environments for enterprise software such as Salesforce, Workday, and email clients. While most test environments settle for stateless API endpoints, Arga builds a full-scale digital twin of your program, effectively replicating your entire enterprise program with permission systems and webhooks intact. The result is a more robust way to train agents across multiple systems.
CEO and co-founder Philip Li gives an example of a prospect generating leads in Salesforce while their colleagues reach out to them individually through Hubspot.
“Can the agent correctly recognize that these two companies are the same company?” Li says, “Can they see if they sent an email only once? Can they identify who to send an email to on two occasions?”
Agent systems still suffer from this kind of ambiguity. And he believes Arga Labs’ tools are essential to improving the system.
Typically, agents can be trained for such tasks through reinforcement learning. So you basically run the scenario tens of thousands of times and only let the successful strategies through. However, the nature of enterprise software makes testing at that scale nearly impossible. There is no easy way to “reset” a system like Salesforce or Outlook, much less replicate it, if you need to run the same scenario again.
Arga Labs’ solution is to create a digital recreation of its software. It replicates its structure, much like a crash test dummy replicates a human being. Arga gives you complete control over your environment, making it easy to reset or change. The company can also run many environments at once to train agents on complex interactions between different programs. The idea is to recreate a person’s complete working environment, with certain tasks overlapping between different programs and knowledge systems.
You can think of this as a way to bridge the reinforcement gap between your coding and other applications. One reason AI coding tools have advanced so quickly is that sophisticated tools already exist to deploy, reverse, and analyze new code. These tools make setting up an RL environment for coding much easier and allow you to test and train AI systems on increasingly complex coding tasks.
These tools don’t yet exist in most business software. But once we do, we can expect AI systems to become much better at using those programs, revolutionizing other industries in the same way they revolutionized coding.
Yuri Sagalov, managing director of General Catalysts, who also runs the company’s seed program, said he sees a growing need for agent testing tools like Arga.
“I think a lot of the economic value you get from agents comes from using business applications,” Sagalov told TechCrunch. “Having a reproducible sandbox environment is very important, much more so for agents than it is for humans.”
If you buy through links in our articles, we may earn a small commission. This does not affect editorial independence.
