Skip to main content
Structured exceptions tell you what failed, whether to retry, and which agent or run was involved — without parsing raw tracebacks.
The user runs the agent; failures raise typed PraisonAIError with category, message, and run context for recovery.

Quick Start

1

Catch any agent error

2

Handle tool errors specifically

3

Catch a failed LLM tool-calling loop


How It Works

Every structured error carries message, agent_id, run_id, error_category, and is_retryable. Subclasses add domain fields such as tool_name or model_name. error_category uses typed kinds such as rate_limit, auth, context_overflow, and billing. LLMResponseError is raised by LLM.get_response() when the tool-calling loop fails and cannot produce a response — previously this was swallowed and returned as an empty string. Catch it to distinguish tool-loop failures; a try/except Exception already covers it. See LLMResponseError.

LLM Response Errors

LLMResponseError is raised when the LLM tool-calling loop hits an exception it cannot recover from. It lives in praisonaiagents.llm alongside the other LLM exceptions:
Behaviour change: a mid-loop tool-calling failure now raises LLMResponseError. Previously the loop swallowed the exception and returned an empty string (""), so agent.start() / agent.chat() looked like they succeeded while silently persisting an empty assistant message and burning retries. Wrap calls that must distinguish a real failure from an empty answer.
LLMResponseError carries a message attribute describing the iteration that failed and chains the original exception via raise … from e, so e.__cause__ holds the underlying error.

Tool Failure Behavior

On the sync path, a tool failure resolves down one of three branches:
  1. Tool raises an exception → the framework wraps it as ToolExecutionError and propagates it.
  2. Tool returns {"error": "..."} → the error dict is handed back to the LLM as a normal tool result, so the model can self-correct (new behavior since PR #4462, confirmed in #4470 which fixes issue #4446. The generic-exception branch in openai_client.py behaves the same way.).
  3. Tool returns a retryable-transient / denial dictToolExecutionError with is_retryable=True, or a short-circuit denial.
Behavior change (PraisonAI PR #4462). A tool that returns {"error": "..."} no longer aborts the run — this was the convention already used by bundled tools like execute_command, tavily_search, and exa_search. Previously the framework raised ToolExecutionError(is_retryable=False) for any dict with a truthy error key, so an empty-command warning from execute_command or a missing-API-key message from a search tool would kill agent.chat() mid-turn. Now the LLM sees the error dict as a normal tool result and can self-correct (retry with a real command, ask the user for the missing key, etc.).Only three cases still raise / short-circuit on the sync path:
  • The tool raised an exception (framework-level failure).
  • The dict is a denial (approval_denied, permission_denied, approval_error, policy_denied, guardrail_denied).
  • The dict is a retryable transient (_outer_timeout on an idempotent tool, circuit_open, or a _praison_retryable marker outside the tool’s retry policy).
Wrap sync chat() / start() calls in try / except ToolExecutionError to catch a raised tool exception:
A tool that returns an error dict (instead of raising) no longer aborts the run — the dict flows back to the model, which retries with corrected arguments:
Behavior change (fix in PraisonAI PR #4252): On the sync path, a tool that raises now raises ToolExecutionError. Previously the sync Agent.chat() silently returned None and, worse, the tool’s error text was treated as an LLM error — so a tool message resembling a rate limit made the framework sleep and re-drive the model with no tool result, then return that hallucinated answer as a success. Callers checking if agent.chat(...) is None: should switch to catching the exception — None cannot distinguish a guardrail rejection from a tool failure anyway.
Sync and async paths are aligned for plain tool-returned {"error": ...} dicts — both hand the error back to the LLM so it can retry with corrected arguments. The two paths still differ for raised exceptions: sync chat() raises ToolExecutionError; async achat() returns an error dict. See Async Tool Safety.

Common Patterns

Retry on transient network failures, fail on config bugs:
Raised errors stop the run; some callbacks record failures on the output instead. See Non-Fatal Errors.

Reaching the Step Limit

When the tool-calling loop reaches ExecutionConfig.max_steps (or max_iter when max_steps is unset), the agent does not hard-cut. On the final permitted step it injects a graceful wrap-up instruction, so the model returns a real summary of what it accomplished and what remains — not a placeholder. Detect truncation with agent.last_stop_reason instead of string-matching:
On a gateway, surface the reason to chat users — see Turn Completion Notes. On the CLI: when a truncated run reaches praisonai run / praisonai-code run, the wrapper reports it distinctly from success — exit code 2 (not 0), status: "truncated" under --output json / --output stream-json with the wrap-up summary preserved in result, plus a one-line stderr warning in interactive mode. See praisonai run → Exit Codes.

Step Budget

Cap tool-use steps and detect graceful truncation with last_stop_reason

Best Practices

Use ToolExecutionError when you only care about tool failures; reserve PraisonAIError for top-level logging.
Include e.error_category, e.agent_id, and e.run_id in observability hooks — they correlate across multi-agent runs.
Validation failures usually mean a programming or config bug. Fix the root cause instead of retrying blindly.
Loop-guard HALT raises ToolExecutionError. Combine with Loop Guard when tools may repeat indefinitely.
A failed tool-calling loop now raises LLMResponseError instead of returning "". Catch it explicitly (from praisonaiagents.llm import LLMResponseError) so retries and observability see the actual error rather than a silent empty string.

Loop Guard

Stop runaway tool loops with HALT/WARN/BLOCK

Non-Fatal Errors

Callback failures captured without crashing