Article

Meet Era: Simulated Companies to Benchmark Models and Eval Agents

Realistic, interconnected data, simulated APIs, and MCP servers for testing, benchmarking, and post-training agents.

Doron Porat
Written by
Doron Porat
Updated on: 
Oct 5, 2026
0
 min read
Meet Era: Simulated Companies to Benchmark Models and Eval Agents

Join 10k Infra & Data Pros

Subscribe to our newsletter for the latest on data protection, analytics, and AI.

Quick Summary

Today, we’re introducing Era, a free tool from Eon that generates simulated companies for benchmarking and improving AI models and agents. Each company comes with realistic, interconnected data, simulated enterprise APIs, and MCP servers your agents can connect to.

Choose an industry, company size, and systems. Select Salesforce, Zendesk, Slack and AWS, for example, and Era creates connected customer records, deals, tickets, conversations, full historical activity and its supporting cloud environment, along with the endpoints, credentials, and MCP connection details needed to use them. Your agent interacts with a simulated company through familiar interfaces that you can swap out to the production environment later on.

For model labs, that means a place to compare checkpoints, find where models struggle, and use graded episodes in your own post-training pipeline. Then evaluate on a separate simulated company to see whether those improvements carry over. 

For agent builders, the same environment supports workflow development, integration testing, model routing, benchmarking and product demos.

Request access · GitHub

The problem: Models and agents can pass benchmarks and still fail in the real world

A model can score well on a benchmark and still struggle to work across a company’s systems. To measure those gaps, train on the failures, and check whether changes help, you need a realistic environment. That requires four things:

  1. Realistic data: Missing fields, duplicate records, uneven volumes, and realistic correlations between customers, deals, and tickets.
  2. Interconnected data with history: The same people, accounts, and events must stay consistent across systems and over time.
  3. Simulated APIs: Salesforce and Zendesk integrations need familiar requests, responses, pagination, and errors.
  4. Simulated MCP servers: Agents need working tools to discover and call against that connected data.

Consider a renewal in Salesforce, a Zendesk ticket marked closed on Monday, and a Thursday Slack message saying the issue returned. Your agent must connect the same customer and issue across all three systems, in the right order. Pristine test data can hide that risk entirely.

Building and maintaining all four is a project of its own, before you even start benchmarking or training. Era provides them together in a simulated company. We built it while developing Eon’s AI features, and now use it for testing, model comparisons, and demos.

The solution: Give models and agents a realistic company to learn and improve in

If you think about it, a company’s customers, employees, and history are spread across tools like Salesforce, Slack, and Zendesk. And since models and agents are not tested against real company data, they tend to fail in the real-world even though they work well during testing and benchmarking. 

Era solves that. Era recreates that setup with realistic, connected data and working interfaces for your agent to use. Each simulated company comes with four components:

1. Realistic data

Era generates missing fields, duplicate records, uneven volumes, and relationships that reflect how companies operate. Customers have different numbers of deals and tickets. Accounts have multiple contacts. Activity varies across teams and over time.

Eon’s experience protecting more than an exabyte of enterprise data informs these patterns. Every record in Era is synthetic.‍

2. Interconnected data with history

Era builds each system’s records from a shared company graph. A customer in Salesforce is the same customer referenced in Zendesk tickets and Slack conversations. Their deals, issues, and interactions follow a shared timeline, so your agent can trace what happened across systems.

3. Simulated APIs

Era exposes that data through vendor-compatible APIs. Select Salesforce or Zendesk, and you get the corresponding simulated endpoints and credentials, with the request and response formats your integration expects. Your agent can query the company through familiar interfaces without you building a mock service for each system.

4. Simulated MCP servers

Era also provides MCP servers with tools your agent can discover and call. Connect using the provided configuration, and your agent can work across the company’s systems against the same interconnected data. You don’t have to build and maintain those MCP servers yourself.

Putting Era and today’s models to the test

We’re working with NVIDIA, Composio, Deel, Decart, Eragon, Openlayer, and Plurai on benchmarks, evaluations, research, and other technical collaborations to make Era more useful for builders. We’ll share more about each partnership over the coming weeks.

As a starting point, our team ran two evaluations to see whether Era could support meaningful evaluations and reveal where models struggle with company work. Its connected records let us ask questions across systems and check responses against computed answers.

Evaluation 1: Finding and connecting records

We tested nine models on 33 questions across Salesforce, Zendesk, and Gong, with three attempts per question and no code execution. Tasks included finding the customer with the largest open pipeline and calculating their total recorded call time.

Accuracy averaged 92.6% on simple filtering, 39.6% on cross-system questions, and 3.7% on multi-hop questions.

Overall accuracy by model. Wrong, FA (false abstention), and NA (no answer) are counts out of 99 attempts per model.

Read the research paper: The Era by Eon Benchmark

Evaluation 2: Finding evidence the obvious record misses

The second evaluation tested whether agents could find and connect evidence beyond the obvious record. For example, Salesforce shows no service credit, but a Gong call contains a promise that must be combined with Zendesk tickets to calculate what is owed.

We tested six models in two agent setups on eight questions, with three attempts per question. The best agent answered 18 of 24 attempts correctly. Four models answered six or fewer in either setup. Adding code execution did not improve the leading models’ results.

Read the research paper: Benchmarking Enterprise Agents on Hidden Knowledge

What you can do with Era

  • Benchmark models and checkpoints on the same enterprise tasks with exact answer keys, so you can compare performance such as cost/token efficiency, tool use efficiency, and understand the failures.
  • Post-train models on graded episodes that target specific weaknesses, using your own training pipeline. Evaluate on a separate company to check whether improvements carry over.
  • Improve agents by comparing prompts, tool choices, and agent frameworks, then rerunning the same tasks to see what helps.
  • Build and test apps and agents across business tools, databases, and cloud storage. Run integration and workflow tests in CI with larger datasets and longer histories.
  • Demo your product and onboard developers using a company shaped like the business you want to show.

Get started

Install the Era CLI: curl -fsSL https://console.era.eon.io/install.sh | sh

Then follow the setup guide to create a company and connect your agent.

# era 0.1.5 | design partner | fintech/mid

# Era
ERA_CONSOLE_URL=https://console.era.eon.io
ERA_TENANT=[REDACTED]
ERA_TENANT_TOKEN=[REDACTED]
ERA_TENANT_TOKEN_EXPIRES_AT=1791598687

# Salesforce
SALESFORCE_BASE_URL=https://salesforce-mcp.era.eon.io
SALESFORCE_MCP_URL=https://salesforce-mcp.era.eon.io/mcp/
SALESFORCE_MCP_HEADER='Authorization: Bearer [REDACTED]'

# Slack
SLACK_BASE_URL=https://slack-mcp.era.eon.io
SLACK_MCP_URL=https://slack-mcp.era.eon.io/mcp/
SLACK_MCP_HEADER='Authorization: Bearer [REDACTED]'

# Zendesk
ZENDESK_BASE_URL=https://zendesk-mcp.era.eon.io
ZENDESK_MCP_URL=https://zendesk-mcp.era.eon.io/mcp/
ZENDESK_MCP_HEADER='Authorization: Bearer [REDACTED]'

‍

The example above shows the .env file that was created for a mid size fintech company with Salesforce, Slack and Zendesk tools inside it. 

How to create benchmarks

To turn this into a benchmark, run each model on the same company state and score its answers against Era’s answer keys. Use the failures to choose post-training tasks, then evaluate the next checkpoint on a separate company. 

See the docs for connection details.

What’s next

We’re adding more systems and industries to broaden the environments available for benchmarking and post-training. We’re also adding questions that ask models and agents to explain why something happened. Era will generate an underlying cause, expose only its indirect traces, and check the inference against the known fact.

Suggest a system or benchmark on Discord or GitHub.

FAQs

Which systems does Era support?

Salesforce, HubSpot, Zendesk, Gong, Jira, Slack, SharePoint, Drive, Eon and Deel, plus AWS services including RDS, S3, EC2, and more. See the docs for available interfaces.

Does Era contain customer data?

No. Every record is synthetic, but the data is generated to reflect how real companies look and operate. Aggregated statistical patterns shape realistic distributions, relationships, and behavior across systems.

Can I use it in CI?

Era's simulators are published as Docker images on Docker Hub, so you can run them as service containers in your pipeline, the same way you'd run Postgres or LocalStack. The GitHub repo includes pytest, Node and GitHub Actions examples. Create an environment for a test run, execute your checks, and tear it down afterward.

Can agents manage their own environments?

Yes. The Era console is also an MCP server, so agents can create and manage environments themselves.

How can I use Era for post-training?

Era provides simulated environments and graded episodes. You can use them to curate demonstrations for supervised fine-tuning or define task-based rewards for reinforcement learning. You supply and run the training pipeline. Keep separate companies for evaluation so you can measure performance beyond the training environments.

FAQ

No items found.
Doron Porat
Doron Porat

Data and AI Strategist