> ## Documentation Index
> Fetch the complete documentation index at: https://docs.get-hive.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Publish and evaluate agents

> Keep agents reliable with evaluation test sets, AI judge, tool and keyword checks, usage analytics, recurring schedules and public agent links.

Once an agent is deployed, your job shifts from building it to keeping it good. Agent Studio gives each agent three places to do that: **Evals** to catch regressions, **Analytics** to see how it's used, and schedules and public links to put it to work.

## Evaluations

The **Evals** tab ("Make sure your agent keeps giving great answers") lets you save **test sets**: example requests paired with what a good answer looks like. Run a test set after changing the instructions, knowledge or model, and you'll see straight away whether the agent got better or worse.

### Build a test set

1. Open the agent and select **Evals**.
2. Click **Create your first test set** (or **New test set** if the agent already has some) and give it a **Test set name**.
3. Choose the checks for the test set, then add cases. Each case has the request, any declared inputs (as a JSON object), the expected answer, and the phrases or tools the checks look for.
4. Click **Create test set** (or **Save test set** when editing).

Conversation shape is **Single turn** (one prompt, one answer). **Multi turn** conversations are marked *coming soon*.

### Checks

Each test set can use any combination of three checks. If none is selected, cases are saved for manual review only:

| Check | What it measures |
| - | - |
| **AI judge** | Compares the agent's answer with the expected answer using a **Judge model** you choose. The default rubric assesses factual agreement using the supplied evidence and does not reward invented facts. |
| **Tool check** | Checks the **Expected tools** against the tools the agent actually used, based on observed evidence rather than the model's claims. |
| **Keyword check** | Checks for **Required phrases** and **Forbidden phrases**, case-insensitive. |

### Run and review

Start a run from the test set. Progress is saved after each case, so a long run can be followed as it goes. **Evaluation results** shows each case's verdict with the observed evidence and usage. **Evaluation history** keeps earlier runs, along with the original agent and test snapshots they ran against, so you can compare revisions fairly.

Deleting a test set keeps its historical results. Roles that can inspect test sets but not create them see "Your role can inspect test sets but cannot create them."

<Tip>
  Add a case every time someone reports a bad answer. Over time your test set becomes a record of every mistake the agent must never make again.
</Tip>

## Analytics

The **Analytics** tab shows how the agent is used:

* **Active users**
* **Conversations**
* **Messages**
* **Messages / User**
* A cost-per-message benchmark
* A **Feedback** section with the helpful and unhelpful ratings people left on answers

Use it to spot agents nobody uses, and agents whose recorded cost per message is higher than you expect. The cost benchmark needs at least 20 messages and doesn't compare agents with each other.

## Schedule an agent

Deployed agents can run on a recurring schedule, for example a Monday-morning pipeline summary or a daily overdue-invoice sweep.

1. Save and deploy the agent. A schedule always runs a deployed revision.
2. Open **Schedule** in the agent editor and choose **Create scheduled task**.
3. Pick a frequency: **Every day**, **Weekdays**, **Every week** or **Every month**. Weekly schedules use **Days of the week**, and monthly schedules use **Day of month** (months without that date are skipped).
4. Set the timezone (for example `Europe/London`) and the prompt or inputs the agent should run with.

Scheduling needs Operator access or above, plus access to both Agent Studio and the Scheduled product. A schedule pins the revision it was created with. If you later deploy a new revision, delete the schedule and create a new one for the new revision.

You can see and manage every schedule in [Scheduled tasks](/library/scheduled-tasks).

## Public agent links

Public links let people outside your workspace talk to an agent without a Hive account, for example a client intake assistant on your website.

<Warning>
  Public agent links are **off by default**, and you can only create links once Hive has turned them on. Talk to us if you'd like them.
</Warning>

When they're available, open **Public links** in the agent editor and choose a saved revision. Each link is tied to that revision, so later edits don't change what the public sees. Library sources you add to a public link go through a scope review. Choosing them doesn't grant permission to publish them.

## Keep improving

<Steps>
  <Step title="Watch feedback and analytics">
    Look for unhelpful ratings and unusual cost per message.
  </Step>

  <Step title="Turn problems into test cases">
    Add the failing request and a good expected answer to a test set.
  </Step>

  <Step title="Edit, re-run evals, then deploy">
    Change the instructions or knowledge, run the test set, and deploy only when results hold up.
  </Step>
</Steps>

## Related

<CardGroup cols={2}>
  <Card title="Build an agent" icon="hammer" href="/agents/build-an-agent">
    Create, test and deploy an agent.
  </Card>

  <Card title="Scheduled tasks" icon="clock" href="/library/scheduled-tasks">
    Manage every recurring agent run in one place.
  </Card>
</CardGroup>
