scanned May 12, 2026

Vellum

vellum.ai

Vellum provides a personal intelligence assistant that helps users manage their tasks and daily life

43/100

Tier 3 · Agent-Accessible

Content answers45/100
Protocol plumbing38/1006 of 16 checks pass

Scored by asking 15 questions a buyer of a ai-ml product asks, then grading this site’s own pages: answered, hedged (partial or vague), or silent (no page answers it). How scoring works

This report is public. Own vellum.ai? Claiming is free: crawl every page, re-audit as you fix, and track your score over time.

Sign in to claim

The fix queue

57 points sit between vellum.ai and 100: 15 open questions and 10 missing protocol checks, ordered by estimated payoff.

Point estimates are per fix under scoring v2. They are not additive to a promised total.

01pricing · importance highGoes silent+7 content pts est.

I'm evaluating Vellum for my side project — what's the exact monthly request limit on the free tier before I hit a hard stop versus just throttling?

What the pages say

No page on the site addresses this.

The fix

Add a dedicated pricing page with clear free tier limits including: exact monthly request/operation count, whether exceeding it triggers hard rejection vs throttling, and behavior when limits are hit.

confidence high · grounding synthesized · weight 0.00 · Absent

02technical · importance highGoes silent+7 content pts est.

I'm building a long-running research workflow that calls external APIs — what's the maximum execution time before Vellum force-terminates a workflow run?

What the pages say

No page on the site addresses this.

The fix

Add explicit documentation of Vellum's server-side workflow execution timeout limit (e.g., 'Workflows are force-terminated after X minutes/hours') to the Long Running Workflows page, likely in a 'Server-Side Limits' or 'Platform Timeouts' section.

confidence high · grounding world-knowledge · weight 0.00 · Absent

03support · importance mediumGoes silent+7 content pts est.

We're on the Team plan and considering Enterprise — what's the actual guaranteed response time for P1 incidents in each tier, not just 'priority support'?

What the pages say

No page on the site addresses this.

The fix

Create a dedicated pricing or support page that explicitly lists guaranteed response times for P1/P2/P3 incidents by plan tier (Team vs Enterprise).

confidence high · grounding world-knowledge · weight 0.00 · Absent

04migration · importance lowGoes silent+7 content pts est.

We're currently on LangChain Cloud with 200+ prompts — does Vellum have any automated import tool for LangChain prompt templates or do we rebuild manually?

What the pages say

No page on the site addresses this.

The fix

Create a dedicated migration guide page that documents whether automated import from LangChain Cloud is available, and if not, provide recommended manual migration steps or contact information for migration assistance.

confidence high · grounding world-knowledge · weight 0.00 · Absent

05technical · importance lowGoes silent+7 content pts est.

I want to build modular workflows — what's the maximum nesting depth for subworkflows, and do deeper nests incur additional latency penalties?

What the pages say

No page on the site addresses this.

The fix

Add documentation on subworkflow nesting limits and performance characteristics, likely in the Subworkflow Node reference docs or a 'Workflow Architecture Limits' page.

confidence high · grounding world-knowledge · weight 0.00 · Absent

Hedged · 10 of 15 questions

A buyer gets something, then has to guess the rest. Expand any row for the evidence and the fix.

06My security team needs to know: is Vellum SOC 2 Type II certified yet, or only Type I? And can I get the report without signing an NDA first?securityHedges+2 pts est.

My security team needs to know: is Vellum SOC 2 Type II certified yet, or only Type I? And can I get the report without signing an NDA first?

What the pages say

Vellum is SOC 2 Type II certified (not just Type I). According to the February 2024 changelog, Vellum achieved SOC 2 Type 2 compliance as of February 28, 2024, and this is confirmed in current documentation stating 'Vellum maintains SOC 2 Type 2 compliance.' However, the site does not specify whether an NDA is required to access the SOC 2 report. To request a copy of the report, Vellum directs users to contact security@vellum.ai or their account representative.

Source: /product/security/data-privacy-and-storage.mdx, /changelog/2024/2024-02.mdx

confidence high · grounding synthesized · weight 0.70 · Answered

07We're hitting rate limits during batch processing — are the API rate limits per-workspace or per-user, and do they scale with your paid tier or stay flat?limitsHedges+2 pts est.

We're hitting rate limits during batch processing — are the API rate limits per-workspace or per-user, and do they scale with your paid tier or stay flat?

What the pages say

The chunks do not explicitly state whether Vellum's API rate limits are per-workspace or per-user. However, they do indicate that limits scale with paid tiers: chunk [7] and [15] mention that 'usage tiers rise automatically with spend' and 'Higher tiers = more TPM/RPM' for OpenAI, and chunk [7] notes that 'enterprise contracts get priority for higher limits' for Anthropic. For Vellum specifically, chunk [8], [9], and [12] reference 'your organization's concurrency limits' when discussing batch processing, suggesting limits are scoped to the organization/workspace level rather than per-user. The chunks do not provide specific details on whether Vellum's own API rate limits scale with paid tiers or remain flat.

The fix

Add explicit documentation on Vellum's API rate limits: clarify whether limits are per-workspace or per-user, and whether limits scale with paid tiers or remain flat across plans. Current content discusses third-party provider limits (OpenAI, Anthropic) and organizational concurrency limits for batch processing, but not Vellum's own API rate limiting structure.

Where we looked: /blog/how-to-manage-openai-rate-limits-as-you-scale-your-app, /blog/how-to-manage-openai-rate-limits-as-you-scale-your-app, /product/workflows/advanced/batching-executions, /product/workflows/advanced/batching-executions

confidence medium · grounding synthesized · weight 0.65 · Vague

08Does Vellum's prompt caching let me configure how long responses stay cached, or is the TTL fixed? I need 24-hour caching for my use case.technicalHedges+2 pts est.

Does Vellum's prompt caching let me configure how long responses stay cached, or is the TTL fixed? I need 24-hour caching for my use case.

What the pages say

Vellum does not let you configure cache TTL — the cache duration is controlled entirely by the model providers (Anthropic and OpenAI), not by Vellum. According to the August 2024 changelog, Anthropic's prompt caching lasts 'up to 5 minutes.' The documentation states that 'Cache durations can vary by model provider' and advises checking provider documentation for specifics. There is no mention of configurable TTL or 24-hour caching options in Vellum's interface.

Source: /changelog/2024/2024-08.mdx, /product/prompts/prompt-caching.mdx, /product/prompts/prompt-caching.mdx, /product/prompts/prompt-caching.mdx

confidence high · grounding synthesized · weight 0.70 · Answered

09I have a fine-tuned model hosted on my own AWS account — can Vellum call it via VPC peering, or does it need to be publicly accessible with an API key?integrationHedges+2 pts est.

I have a fine-tuned model hosted on my own AWS account — can Vellum call it via VPC peering, or does it need to be publicly accessible with an API key?

What the pages say

Vellum supports calling fine-tuned models via API key authentication for specific hosted platforms, but does not explicitly document VPC peering for custom AWS-hosted models. According to the docs, Vellum currently supports: OpenAI models hosted on Azure, OpenAI fine-tuned models, and fine-tuned models hosted on Fireworks AI. These are configured as 'private Model Templates' in Vellum. For models requiring API keys, you provide your own API key through the provider-level settings in the Models page. The documentation mentions that 'some like those from AWS Bedrock will require some additional steps taken within your account,' but does not specify whether VPC peering is supported for custom AWS-hosted fine-tuned models. Vellum does support self-hosted/VPC deployments of Vellum itself, but this refers to running the Vellum platform in your infrastructure, not connecting to your custom model via VPC peering.

The fix

Add explicit documentation on whether Vellum supports VPC peering, PrivateLink, or other private connectivity options for custom models hosted in customer AWS accounts, or clearly state that such models must be publicly accessible with API key authentication.

Where we looked: /product/prompts/custom-models.mdx, /product/prompts/custom-models, /developers/workflows-sdk/configuration.mdx, /developers/client-sdk/introduction, /developers/client-sdk/introduction.mdx

confidence medium · grounding synthesized · weight 0.65 · Vague

10For the Code Execution Node, what's the memory limit and can I install arbitrary Python packages, or am I restricted to a pre-approved list?technicalHedges+2 pts est.

For the Code Execution Node, what's the memory limit and can I install arbitrary Python packages, or am I restricted to a pre-approved list?

What the pages say

For the Code Execution Node, you can install arbitrary Python packages by specifying them in the `packages` attribute with exact versions (e.g., `CodeExecutionPackage(name="pandas", version="2.0.0")`). The chunks do not specify a memory limit for Code Execution Nodes. However, chunks [10], [15], and [18] reference "usage limits" for Container Images at a linked URL (https://docs.vellum.ai/changelog/2025/2025-07.mdx/developers/workflows-sdk/custom-container-images#usage-limits), but that specific page content was not retrieved.

The fix

Add explicit documentation of memory limits (e.g., '512 MB RAM per execution') directly on the Code Execution Node documentation page, not just via a link to a separate 'custom container images' page.

Where we looked: /developers/workflows-sdk/api-reference/nodes/code-execution-node, /developers/workflows-sdk/api-reference/nodes/code-execution-node, /product/workflows/nodes/code-execution-node.mdx, /changelog/2025/2025-07.mdx, /changelog/2025/2025-07.mdx

confidence medium · grounding synthesized · weight 0.65 · Vague

11Our team needs dedicated infrastructure for compliance reasons — at what user or workload threshold does Vellum typically offer custom VPC deployment versus shared cloud?pricingHedges+2 pts est.

Our team needs dedicated infrastructure for compliance reasons — at what user or workload threshold does Vellum typically offer custom VPC deployment versus shared cloud?

What the pages say

Vellum offers VPC deployment for "enterprise customers with specific security requirements" but does not specify a user or workload threshold for when custom VPC deployment is offered versus shared cloud. The chunks indicate VPC is available as an enterprise option for compliance and security needs, not based on scale metrics. For VPC solutions, Vellum directs customers to "contact your team's VPC / infrastructure administrator" or "reach out to Vellum support."

The fix

Add explicit criteria to the pricing or enterprise page stating the user count, workload volume, or spend threshold that triggers VPC eligibility versus shared cloud, or clarify if VPC is purely security-driven rather than scale-driven.

Where we looked: /product/security/static-ips.mdx, /product/security/static-ips, /blog/top-langchain-alternatives, /blog/top-langchain-alternatives, /blog/top-langchain-alternatives, /blog/top-langchain-alternatives

confidence medium · grounding synthesized · weight 0.65 · Vague

12If I delete a prompt or workflow, how long does Vellum retain the underlying execution logs and what regions is that data stored in?securityHedges+2 pts est.

If I delete a prompt or workflow, how long does Vellum retain the underlying execution logs and what regions is that data stored in?

What the pages say

By default, interaction/execution data is stored indefinitely. Enterprise customers can configure data retention policies to automatically delete monitoring data after 30, 60, 90, or 365 days. However, the chunks do not specify what happens to execution logs when a prompt or workflow is deleted (vs. automatic policy-based deletion), nor do they specify which regions this data is stored in.

The fix

Add explicit documentation on: (1) whether deleting a prompt/workflow immediately deletes its execution history or if retention policies still apply, and (2) which geographic regions Vellum stores execution data in.

Where we looked: /product/security/data-privacy-and-storage.mdx, /product/organizations/data-retention-policies.mdx, /product/organizations/data-retention-policies.mdx, /product/organizations/data-retention-policies.mdx

confidence medium · grounding synthesized · weight 0.65 · Vague

13In the Agent Builder, how many tools can a single agent invoke in one turn before hitting a limit, and can that be configured?technicalHedges+2 pts est.

In the Agent Builder, how many tools can a single agent invoke in one turn before hitting a limit, and can that be configured?

What the pages say

The Agent Node has a configurable "Max Prompt Iterations" setting that controls the maximum number of times the model can call tools before the node terminates execution. However, the chunks do not specify a hard limit on how many tools can be invoked in a single turn, nor whether this is specifically configurable within the Agent Builder interface versus the Agent Node configuration.

The fix

Add explicit documentation about whether Agent Builder exposes or configures tool invocation limits, or clarify that limits are set at the Agent Node level after workflow generation.

Where we looked: /product/workflows/nodes/agent-node.mdx, /product/workflows/nodes/agent-node.mdx, /changelog/2025/2025-10.mdx, /changelog/2025/2025-10

confidence medium · grounding synthesized · weight 0.65 · Vague

14Does the Guardrail Node come with pre-built content safety classifiers, or do I need to bring my own moderation API for PII and toxicity detection?technicalHedges+2 pts est.

Does the Guardrail Node come with pre-built content safety classifiers, or do I need to bring my own moderation API for PII and toxicity detection?

What the pages say

The Guardrail Node executes pre-defined Metrics that are created in the Vellum UI, but the chunks do not specify whether Vellum provides pre-built content safety classifiers for PII and toxicity detection, or if users must bring their own moderation API. The documentation shows examples of custom metrics (like 'response_quality_metric' with 'safety' as an evaluation criteria) and references to external metrics like Ragas Faithfulness, but does not explicitly address built-in PII/toxicity classifiers.

The fix

Add explicit documentation about what built-in content safety classifiers (PII detection, toxicity detection, etc.) are available out-of-the-box versus what requires custom configuration or external API integration.

Where we looked: /developers/workflows-sdk/api-reference/nodes/guardrail-node.mdx, /product/workflows/nodes/guardrail-node.mdx, /developers/workflows-sdk/api-reference/nodes/guardrail-node.mdx

confidence medium · grounding synthesized · weight 0.65 · Vague

15For the image and document inputs in prompts, what's the maximum file size and are there page limits for PDFs specifically?limitsHedges+1 pt est.

For the image and document inputs in prompts, what's the maximum file size and are there page limits for PDFs specifically?

What the pages say

For PDFs used as prompt inputs: maximum file size is 32MB. For images: maximum file size is 32MB (with combined images not exceeding this limit for multiple images). No page limits for PDFs are mentioned in the documentation.

Source: /product/prompts/multimodality.mdx, /product/prompts/multimodality.mdx

confidence high · grounding stated · weight 0.80 · Answered

Protocol plumbing · 38/1006 of 16 checks pass · each fix +6 protocol pts est.

The other half of the score: 16 checks for the files and headers agents look for. The 10 below are installs, not judgment calls, and most are an afternoon. Expand any for the snippet and the standard it follows. They sit after the queue because none of them changes what your pages say.

llms.txtDiscoverability+6 pts est.

Sitedex generates this file from your crawl. Grab it in Files from this audit below.

Standardllmstxt.orgCommunity spec

Content signalAccess+6 pts est.
Install snippet
User-agent: *
Content-Signal: search=yes, ai-input=yes, ai-train=no
Allow: /

StandardCloudflare proposalVendor proposal

Clean crawlAccess+6 pts est.

StandardSitedex metricSitedex metric

Markdown negotiationRendering+6 pts est.

StandardRFC 9110 + 7763IETF RFC

MCP cardInteraction+6 pts est.

Sitedex generates this file from your crawl. Grab it in Files from this audit below.

StandardModel Context ProtocolCommunity spec

OpenAPI specInteraction+6 pts est.

StandardOpenAPI SpecIndustry standard

WebMCP widgetInteraction+6 pts est.

Sitedex generates this file from your crawl. Grab it in Files from this audit below.

StandardW3C WebMCP draftW3C / WHATWG

Canonical URLsHygiene+6 pts est.
Install snippet
<link rel="canonical" href="https://vellum.ai/" />

StandardRFC 6596IETF RFC

Organization schemaIdentity+6 pts est.

Sitedex generates this file from your crawl. Grab it in Files from this audit below.

StandardSchema.org + JSON-LDIndustry standard

Sitemap lastmodDiscoverability+6 pts est.

Standardsitemaps.orgIndustry standard

Already passing 6 of 16: robots.txt, sitemap.xml, AI crawler access, Server-rendered content, Meta descriptions, HTML lang attribute.

Ask this site’s index

Sitedex already serves vellum.ai as an MCP endpoint. Ask vellum.ai anything an AI agent might ask, and see what its index returns. (To score your own site, use the form below.)

Snippets & configs

For developers and the engineer-on-call: copy these into your tools or your site.

Files from this audit

Built from this crawl. Download or copy each, then install it at the path noted.

llms.txt

Built from this crawl. Install at /llms.txt so agents start here.

organization.json

Organization JSON-LD, pre-filled from this crawl. Wrap in a ld+json script.

server-card.json

MCP server card built from this crawl. Host at /.well-known/mcp/server-card.json.

webmcp.json

WebMCP discovery manifest built from this crawl. Host at /.well-known/webmcp.json.

MCP endpoint

https://mcp.sitedex.dev/s/vellum-ai/mcp

The URL anyone's agent points at. Read-only; safe to share.

Claude Code

claude mcp add vellum --transport http https://mcp.sitedex.dev/s/vellum-ai/mcp

One command, then the agent has it.

Cursor / Continue

{
  "mcpServers": {
    "vellum": {
      "url": "https://mcp.sitedex.dev/s/vellum-ai/mcp"
    }
  }
}

Drop into mcp.json.

WebMCP: two parts

WebMCP-capable browsers run the widget at runtime. Crawlers without JS rendering need the discovery manifest to find your tool surface. Install both.

1 · Widget script

<script async src="https://sitedex.dev/widget.js"></script>

Drop in <head>. WebMCP-capable browsers (Chrome 146+ Origin Trial) call navigator.modelContext.provideContext() via this script.

2 · Discovery manifest

{
  "$schema": "https://wellknownmcp.org/schemas/webmcp.json",
  "name": "vellum.ai",
  "tools": [
    { "name": "search", "description": "Search vellum.ai's indexed content." },
    { "name": "get_page", "description": "Fetch a page from vellum.ai as markdown." }
  ]
}

Host alongside the script at /.well-known/webmcp.json. Crawlers that don't render JS rely on this.

Your turn

See which of these questions your site goes silent on.

Free, about 5 minutes. We crawl your site, test it against the buyer questions your category asks, and name what’s vague, contradictory, or missing, plus the files AI agents look for.

ComingEmbeddable grade badgeScore history and deltasOpt-in public board