AWS Reference
Commands, customization, environment variables, and troubleshooting for the AWS template.
The template names the Express service agent-os, the ECR repo agentos, and the RDS instance agentos-db. Secrets live under agentos/* in Secrets Manager, and the scripts record the service ARN and region in tmp/agentos-aws.state.
Manage
| Task | Command |
|---|---|
| Deploy code changes | ./scripts/aws/redeploy.sh |
| Sync env variables | ./scripts/aws/env-sync.sh (prefers .env.production, falls back to .env; pass a file explicitly to select it) |
| Tail logs | aws logs tail /ecs/agent-os --follow --region <region> |
| Watch a rollout | aws ecs monitor-express-gateway-service --region <region> --service-arn <arn> |
| Tear down | ./scripts/aws/down.sh (add --yes to skip the confirmation) |
Production auth
Token-Based Authorization is on by default. Production startup requires JWT_VERIFICATION_KEY or a readable JWKS file at the container path in JWT_JWKS_FILE; otherwise the process exits.
Token-Based Auth gives you three things:
- Protected runtime access. Protected API routes require a valid credential. Health, discovery, and API documentation remain public; Slack and MCP OAuth have their own authentication flows.
- Per-request identity. Middleware validates the token and exposes its
user_id, optionalsession_id, scopes, and claims to the request. - Scope-based permissions. Token scopes control access to AgentOS routes and resources.
The template already sets AuthorizationConfig(user_isolation=True). Authenticated non-admin REST access is scoped to the principal; local dev mode disables scope enforcement and is open when no credentials are configured. This does not scope Platform Manager’s direct database tools to that REST identity. See User Isolation for the boundaries and admin exceptions.
/health stays open. AgentOS serves it unauthenticated even in production, so the ALB health checks pass.
To disable JWT authentication in a private deployment with another auth layer, set authorization=False, remove JWT_VERIFICATION_KEY and JWT_JWKS_FILE from the running service, and rebuild or redeploy the app as needed. Disabling scope enforcement alone does not remove JWT validation while those credentials remain configured. MCP OAuth remains enabled while MCP_CONNECT_SECRET is set.
Customize
Ask your coding agent to run /create-agent, or do it by hand. Create agents/my_agent.py:
from agno.agent import Agent
from app.learning import shared_learning
from app.settings import default_model
from db import get_postgres_db
INSTRUCTIONS = """\
What the agent does, which tools it uses, the rules to follow when answering.
"""
my_agent = Agent(
id="my-agent",
name="My Agent",
user_id="anonymous-user", # Local fallback; authenticated runs supply identity.
model=default_model(),
db=get_postgres_db(),
instructions=INSTRUCTIONS,
learning=shared_learning,
add_datetime_to_context=True,
add_history_to_context=True,
num_history_runs=5,
)The fallback user ID lets local anonymous calls use the shared learning machine; those calls share one profile. Authenticated run identity overrides this default.
Import it in app/main.py and update the agents argument in the existing AgentOS call. Keep its other arguments, including teams, workflows, knowledge, and registry:
from agents.my_agent import my_agent
agent_os = AgentOS(
# Keep the other arguments from the existing call.
agents=[platform_builder, platform_manager, platform_engineer, my_agent],
)Add its UI metadata beneath the existing manifest: key in app/config.yaml:
my-agent:
description: "What the agent does."
quick_prompts:
- "First example prompt"
- "Second example prompt"
- "Third example prompt"Local containers reload Python source changes. After editing app/config.yaml, run docker compose restart agentos-api to reload the manifest. For production, run ./scripts/aws/redeploy.sh.
app/settings.py defines default_model(), used by every agent. Change it in one place:
from agno.models.anthropic import Claude
def default_model():
return Claude(id="claude-sonnet-5")Add anthropic to pyproject.toml, set the provider key in your env, and regenerate pins:
./scripts/generate_requirements.shRebuild locally with docker compose up -d --build. For production:
./scripts/aws/redeploy.shIt rebuilds the image and syncs .env.production in one pass.
Agno ships 100+ toolkits. See Toolkits.
from agno.tools.slack import SlackTools
my_agent = Agent(
# Keep the agent’s existing configuration.
tools=[SlackTools()],
)- Edit
pyproject.toml. - Regenerate pins:
./scripts/generate_requirements.sh(addupgradeto refresh every pin). - Rebuild locally with
docker compose up -d --build, or redeploy with./scripts/aws/redeploy.sh.
Set both variables in your env file:
SLACK_BOT_TOKEN=xoxb-...
SLACK_SIGNING_SECRET=...Sync with ./scripts/aws/env-sync.sh; both land in Secrets Manager. The interface activates automatically and routes messages to the Agno team. Change team= in app/main.py to select another team, or replace it with agent=my_agent to route to an agent. See Slack setup.
The deployment check runs daily by default (ENABLE_DEPLOY_CHECK=True); it uses fixed checks without model calls. The run-evals schedule is always registered but starts disabled because it uses model calls. Enable it from the AgentOS UI. Both workflows remain runnable on demand. Startup reapplies ENABLE_DEPLOY_CHECK to the deployment-check schedule; the enabled state of an existing run-evals schedule is preserved.
ARM cuts the Fargate line item from about $70 to about $57 per month. Edit runtimePlatform in scripts/aws/task-def.json, change docker build --platform linux/amd64 to linux/arm64 in both up.sh and redeploy.sh, then run ./scripts/aws/redeploy.sh.
Format, validate, and run evals
Run evals against a dedicated local test platform with no concurrent writers. The starter’s cleanup hooks remove components and learning state created during a case; concurrent application writes can be removed too. The same prerequisite applies to scheduled evals. See eval setup and isolation.
The host scripts require uv. The setup script creates a Python 3.14 venv:
./scripts/venv_setup.sh
source .venv/bin/activate| Task | Command |
|---|---|
| Format | ./scripts/format.sh |
| Lint and type-check | ./scripts/validate.sh |
| Run smoke evals | python -m evals --tag smoke |
./scripts/mcp_check.sh runs inside the container, so it needs no venv.
Environment variables
| Variable | Required | Default | Description |
|---|---|---|---|
OPENAI_API_KEY | Yes | - | Models and embeddings. |
RUNTIME_ENV | No | prd | dev disables scope enforcement; configured JWT credentials still enable token validation. Compose sets it for local. Keep production on prd so scope enforcement remains enabled. |
JWT_VERIFICATION_KEY | Production | - | Public key from os.agno.com. Quote the value so the multi-line PEM parses as one variable. |
JWT_JWKS_FILE | Production | - | Path inside the running container to a JWKS JSON file. The scripts set only this path. Add the file to the image build context, rebuild, and redeploy the image, or configure a platform mount and roll the service. |
MCP_CONNECT_SECRET | No | generated by up.sh | OAuth consent secret (16+ chars) for connecting claude.ai and ChatGPT to /mcp. up.sh generates one on deploy and writes it to .env.production. |
AGENTOS_MCP_SIGNING_KEY | No | generated | Optional high-entropy signing-key material (32+ chars) for OAuth tokens. Unset, a strong key is generated and persisted in the database. Rotating it invalidates outstanding tokens. |
AGENTOS_URL | No | http://127.0.0.1:8000 | Scheduler base URL. up.sh sets it to your Express service URL. Loopback reaches the app inside this container; set the public URL for hosted MCP OAuth and the template’s deployment check. When MCP_CONNECT_SECRET is set, OAuth metadata also derives its public origin from this URL. |
SERVICE_ARN | No | written by up.sh | Deploy metadata the scripts use to find the Express service. env-sync.sh never sends it to the container. |
AWS_REGION | No | us-east-1 | AWS region used for provisioning and lifecycle commands. |
ENABLE_DEPLOY_CHECK | No | True | Daily deployment-check cron. |
EVALS_TAG | No | smoke | Eval tag the run-evals workflow runs. |
EVALS_CASE_TIMEOUT_SECONDS | No | 90 | Fallback timeout for cases without an explicit timeout. |
EVALS_SUITE_TIMEOUT_SECONDS | No | derived from selected cases | Sum of selected case timeouts plus 30 seconds per case, with a 60-second floor. A positive integer overrides this ceiling. |
PARALLEL_API_KEY | No | - | WebSearch uses the Parallel SDK when set, keyless MCP otherwise. |
SLACK_BOT_TOKEN | No | - | Set with the signing secret to enable Slack. |
SLACK_SIGNING_SECRET | No | - | Set with the bot token to enable Slack. |
DB_HOST / DB_PORT / DB_USER / DB_DATABASE | No | matches compose | Postgres connection. env-sync.sh skips these; production values come from the provisioned RDS instance. |
DB_PASS | No | matches compose | Postgres password. up.sh generates the production value and stores it in Secrets Manager. |
DB_DRIVER | No | postgresql+psycopg | SQLAlchemy driver. env-sync.sh skips it too; production uses the task-definition value. |
AGNO_DEBUG | No | False | Verbose Agno logs. Compose sets it for dev. |
WAIT_FOR_DB | No | False | If True, the entrypoint blocks on the database before starting. Compose sets it. |
Troubleshooting
Upgrade the AWS CLI (for example brew upgrade awscli) until aws ecs create-express-gateway-service help works. If credentials are the problem instead, run aws configure and confirm aws sts get-caller-identity succeeds.
The RDS instance deploys into the region's default VPC. Create one with aws ec2 create-default-vpc, or adapt scripts/aws/up.sh to your own VPC.
Expected. At os.agno.com, choose Connect OS → Live, enter your service URL, name it Live AgentOS, turn on Token-Based Authorization (JWT) on the connection panel, and connect. The UI generates the public key. If the OS is already connected, enable the setting under Settings → OS & Security. Paste the full PEM into the script prompt. To add a PEM later, set JWT_VERIFICATION_KEY and run ./scripts/aws/env-sync.sh. To use JWKS, add the file to the image build context and rebuild, or configure a mount. Set JWT_JWKS_FILE to its container path, then redeploy or roll the service. Env sync alone only updates the path.
Non-dev mode requires JWT verification configuration. Configured JWT credentials also enable validation in dev mode. Set JWT_VERIFICATION_KEY and sync. For JWKS, verify the file exists inside the container at JWT_JWKS_FILE; changing the variable alone does not deliver it. For a private deployment using another auth layer, follow the credential-removal steps under Production auth.
First-time provisioning of the ALB, certificate, and DNS takes 10-25 minutes; up.sh waits through it. Past that window, the known first-run cause is freshly created IAM roles: Express's async infrastructure calls get denied before the role policies propagate, and ECS never retries. up.sh detects this and recreates the service once; the second attempt provisions reliably. If it still stalls, inspect with aws ecs monitor-express-gateway-service --region <region> --service-arn <arn>, look for an AccessDenied CreateLoadBalancer event in CloudTrail, then delete the service and re-run ./scripts/aws/up.sh.
The gateway is up; the app is still starting. First boot pulls the image and waits for the database. Wait a few minutes and check aws logs tail /ecs/agent-os --follow --region <region>.
Check that the schedule is enabled, the scheduler is running, and its request URL is reachable from the container. Inspect application logs for request or authentication failures. The default loopback URL reaches the app on port 8000; it does not by itself prevent scheduled runs. For hosted MCP OAuth, set a public AGENTOS_URL and run ./scripts/aws/env-sync.sh.
The scripts resolve the service ARN from tmp/agentos-aws.state first, then from a SERVICE_ARN= line in .env.production or .env. On a fresh clone or a new machine, add SERVICE_ARN=arn:aws:ecs:... with your service’s actual ARN to .env.production. Set AWS_REGION in your shell to the service’s region.
The commands are hitting the wrong region. The scripts use AWS_REGION if set, then the region recorded in tmp/agentos-aws.state, then us-east-1. Set AWS_REGION to the region you deployed to and re-run ./scripts/aws/down.sh.