AI Employees — Discovery & Analysis
Status: Exploration / Discovery Audience: Founder, operators Scope: Internal business operations (marketing, content, social, CX, SEO, email lifecycle, analytics). Not a consumer-facing feature for Objectuve users — the existing AI Coach covers that surface. Last updated: 2026-06-04
1. Executive Summary
The question: Should Objectuve adopt a Sintra-AI-style "AI Employees" platform to run business operations, or build something equivalent in-house?
The answer (short version): Hybrid. Buy a no-code agent builder (Gumloop or Zapier Agents) for low-stakes, integration-heavy operations — social scheduling, lead routing, analytics digests, inbox triage. Build a custom "AI Workforce" service on top of the already-planned LiteLLM proxy for brand-critical roles where Objectuve's existing Claude Code skill library is a genuine moat — content writing, lifecycle email drafting, customer support, SEO audits.
Skip Sintra. Its helpers are powered by commodity LLMs (GPT-4.1 + Claude 4 Sonnet) wrapped in named personas, and the 250-credits/month ceiling caps all tiers. The underlying models are already directly accessible through the LiteLLM proxy Objectuve is building. Paying $97/mo for persona wrappers doesn't pencil when you have a richer, more brand-aligned prompt library already in .claude/skills/.
The key insight: Objectuve isn't starting from zero. The Claude Code harness already contains ~40 specialized skills that each read like a job description — crafting-page-messaging, designing-lifecycle-messages, orchestrating-social-rhythm, inspecting-search-coverage, tightening-brand-voice, and many more. What's missing isn't domain expertise. It's autonomous execution, scheduling, memory, and oversight.
2. Context & Motivation
Objectuve is operated by a lean team and sits at the intersection of several specialized disciplines: product, brand, marketing, content, social, SEO, analytics, and customer support. Each discipline eats time that doesn't compound the product itself. The promise of "AI Employees" is that a small set of always-on digital workers can absorb the repetitive parts of those disciplines.
This doc exists to scan the landscape before committing to anything. No vendor has been selected. No code has been written. The goal is to leave the reader with enough context to pick a direction and run a concrete experiment.
Two pieces of existing context shape every recommendation below:
Objectuve already has a specialized skill library. The Claude Code harness contains ~40 skills covering nearly every business operation a solo-to-small team would want to automate. Each skill encodes Objectuve-specific constraints — brand voice, anti-social philosophy, stack conventions, legal entity details. This is an asset no vendor can replicate out of the box.
A LiteLLM proxy extraction is already planned. See
docs/product/completed/dedicated-ai-service-prd.md. That service will give Objectuve per-user/per-feature cost tracking, model routing, fallbacks, caching, and prompt versioning. It's the governance layer any serious AI workforce needs, and it's already on the roadmap for other reasons.
Those two facts collapse the build-vs-buy tradeoff. Building is not a greenfield project — it's stitching existing pieces together.
3. What Are "AI Employees"?
The category is new and the terminology is muddy. Four things often get conflated:
| Category | What it is | Example | Fit for Objectuve |
|---|---|---|---|
| Chatbots | Reactive Q&A, single-turn, no memory | ChatGPT, Claude.ai | Already have via Claude Code |
| Copilots | Suggest-and-wait assistants embedded in a workflow | GitHub Copilot, Intercom Fin | Not what we're discussing |
| Agent frameworks | Libraries for building multi-step autonomous agents | CrewAI, LangGraph, Claude Agent SDK | Build-path foundation |
| AI Employees | Role-based autonomous agents with persistent identity, memory, scheduled execution, and integrations into existing tools | Sintra, Lindy, Arahi | The category we're evaluating |
The distinguishing feature of an AI Employee is autonomy plus persistence. It has a role ("Social Media Manager"), a schedule ("posts 3x/week"), memory ("last month's engagement beat Wednesday posts"), and the ability to execute without waiting for a prompt. That's what Sintra markets and what Lindy, Gumloop, and Arahi all implement differently.
4. Sintra AI Deep Dive
Sintra is the most visible product in the "AI Employees" category and the reason this doc exists.
What you get
12 named helpers, each with a persona and a domain. The marquee roster:
- Cassie — Customer Support Specialist
- Penn — Copywriter
- Soshie — Social Media Manager
- Seomi — SEO Specialist
- Milli — Sales
- Buddy — Business Development Manager
- …plus six more covering data analysis, email, design briefs, research, etc.
What's under the hood
Sintra is a thin orchestration layer over commodity LLMs. Third-party reviews confirm the helpers are powered by GPT-4.1 and Claude 4 Sonnet, with image generation via Gemini 2.5 Flash and Flux Kontext. In other words, Sintra's moat is the persona/prompt layer and the UI — not the underlying intelligence. (Sources: Social Rails review, Zaturn AI review.)
Pricing
Plans run $39–$97/month. The flagship tier (Sintra X, $97/mo) unlocks all 12 helpers plus "Brain AI" and "Power-Ups." Heavy promotional discounting is common for first-time buyers.
The ceiling that matters: all plans share a 250 credits/month cap. Moderate-to-heavy users run out before month-end regardless of tier. That's the biggest published complaint on third-party reviews.
Strengths
- Fast onboarding — no configuration beyond signing up and picking a helper
- Specialized personas feel more focused than generic ChatGPT prompts
- 4-star Trustpilot rating across ~8,000 reviews — real users are getting real value
- Good fit for solo operators who don't already have a prompt library
Weaknesses
- The execution gap. Reviews consistently note Sintra produces deliverables but leaves implementation to the user. You still have to copy-paste the social post into Buffer, the email into Mailtrap, the SEO changes into your CMS. Sintra writes; you execute.
- The credit ceiling. 250/month is fine for exploration, punishing for daily use.
- No brand-voice customization beyond prompts. The personas have fixed tones. Tightening to Objectuve's "coach who's also a friend" voice requires prompt engineering on every interaction.
- No integrations to Objectuve's existing stack. PostHog, Sentry, Mailtrap, Clerk, Cloud Run — none of them plug in natively.
Who Sintra is actually for
Solopreneurs who (a) don't already have a prompt library, (b) aren't picky about brand voice, (c) need to generate drafts more than execute actions, and (d) are comfortable living inside Sintra's UI rather than their existing tools. Objectuve checks none of those boxes.
5. Alternatives Landscape
Grouped by category, because the tradeoffs differ sharply across tiers.
Managed "AI Employee" products
| Product | Positioning | Pricing | Notable |
|---|---|---|---|
| Sintra | 12 named helpers, persona-first | $39–$97/mo, 250 credits | Execution gap; thin wrapper over GPT-4.1/Claude |
| Arahi | Autonomous execution vs suggest-and-wait; 1,000+ integrations | Enterprise-oriented | Closer to true autonomy than Sintra |
| 11x.ai | Vertical: SDR/outbound sales | $$$ | Not relevant — wrong vertical |
| Artisan.co | Vertical: outbound sales | $$$ | Not relevant — wrong vertical |
No-code agent builders
| Product | Positioning | Pricing | Notable |
|---|---|---|---|
| Lindy | General-purpose agent platform, 4,000+ integrations | From ~$49/mo | Strong template library, email/calendar strength |
| Gumloop | Visual workflow automation for AI | From ~$37/mo | Used by Canva, Autodesk, Rakuten. Best for data-heavy pipelines. |
| Relevance AI | Data analysis and research automation | $19–$234/mo, usage-based | Usage-based pricing can get expensive at scale |
| Zapier Agents | Natural-language agents across 8,000+ apps | Tied to Zapier plan | Widest integration surface; weakest reasoning |
| n8n | Open-source workflow automation with AI nodes | Self-host free or cloud $20+ | Hackable; good for technical operators |
DIY multi-agent frameworks
| Framework | Positioning | Notable |
|---|---|---|
| Claude Agent SDK | Tool-use-first, MCP-native, agents-invoke-agents | Same mental model Claude Code uses |
| CrewAI | Role-based DSL, lowest learning curve | 44k+ GitHub stars, native MCP + A2A |
| LangGraph | Stateful graph orchestration with checkpointing | Best for complex branching/loops |
| OpenAI Agents SDK | Fastest path to "something working" | Opinionated toward OpenAI models |
Sources: Gumloop's Sintra alternatives roundup, Lindy's Sintra alternatives roundup, G2's Sintra competitors, Agent framework comparison 2026.
Comparison matrix
| Dimension | Managed (Sintra) | No-code (Lindy/Gumloop) | DIY (Claude Agent SDK) |
|---|---|---|---|
| Time to first value | Hours | Days | Weeks |
| Autonomy ceiling | Low (draft + copy) | Medium (scheduled workflows) | High (full agent loops) |
| Brand-voice control | Low | Medium | High |
| Integration with Objectuve stack | None | Some (via Zapier/webhooks) | Full |
Reuse of .claude/skills/ library | None | None | Full |
| Pricing model | Flat + credits | Flat or usage | Token + infra cost |
| Oversight burden | Low | Medium | High |
| Who owns the IP | Vendor | Vendor workflows, your data | You |
6. What Objectuve Actually Needs
Mapping concrete business operations to roles, and for each role checking whether the existing Claude Code skill library already covers the "how" — meaning the only thing a custom build would need to add is scheduling, memory, and integration.
| Role | Business need | Existing skill(s) that cover it | Autonomous-ready? |
|---|---|---|---|
| Content writer | Marketing landing copy, blog posts, PRD drafts | crafting-page-messaging, framing-release-stories, tightening-brand-voice | Yes, with approval queue |
| Lifecycle email designer | Welcome series, re-engagement, milestone emails | designing-lifecycle-messages, tightening-brand-voice | Yes, with approval queue |
| Social media manager | Cadence planning, post drafts, channel mapping | orchestrating-social-rhythm, planning-editorial-arcs | Partial — needs platform API integrations |
| SEO specialist | Meta tags, Schema.org, sitemap audits, keyword research | inspecting-search-coverage, adding-structured-signals, scaling-template-pages | Yes — mostly file edits on the repo |
| Growth experimenter | Hypothesis generation, A/B test design, funnel analysis | generating-growth-hypotheses, designing-variation-tests, running-product-experiments, mapping-conversion-events | Partial — needs PostHog read access |
| Customer support | First-response drafts, FAQ updates, tone enforcement | refining-prompt-surfaces, tightening-brand-voice | Partial — high risk, needs approval gates |
| Paid campaign planner | Ad creative, UTM strategy, landing page alignment | calibrating-paid-campaigns, crafting-page-messaging | Partial — needs ad platform integrations |
| Onboarding auditor | Funnel drop-off review, wizard copy refinement | accelerating-first-run, designing-onboarding-paths, streamlining-signup-steps | Yes — reads PostHog, edits Vue files |
| Release-notes writer | Changelog drafts, release narratives | framing-release-stories, writing-release-notes | Yes, with approval queue |
| Analytics reporter | Weekly funnel digest, activation metrics review | instrumenting-product-metrics, mapping-conversion-events | Yes — needs PostHog MCP |
The pattern: for 7 out of 10 roles, the Claude Code skill library already encodes the domain expertise. Objectuve's problem is not "what should the AI know?" — it's "how does the AI run on a schedule, remember what it did last week, and ship the output to the right place?"
That's exactly what agent frameworks solve.
7. Pros & Cons by Category
Managed "AI Employee" products (Sintra-style)
Pros
- Zero setup cost, instant onboarding
- Good UX for non-technical operators
- Updated by the vendor — no maintenance burden
Cons
- Execution gap — they produce, you deploy
- No integration with Objectuve's stack
- Brand voice is generic or prompt-injected
- Credit ceilings cap serious daily use
- Vendor lock-in; can't reuse
.claude/skills/library - Paying for thin prompt layers on top of models you already access via LiteLLM
No-code agent builders (Lindy, Gumloop, Zapier Agents)
Pros
- Real autonomous execution (not just drafts)
- Huge integration surfaces (4,000–8,000 apps)
- Visual workflows — easy to audit and debug
- Can ship low-stakes automations in an afternoon
- Good middle ground for operators who want more than Sintra but aren't ready to write code
Cons
- Still can't reuse Objectuve's custom skill library
- Brand-voice control is limited to prompt injection
- Usage-based pricing can surprise you at scale
- Vendor lock-in on the workflow layer (your logic lives in their DSL)
- Not well-suited for multi-step reasoning that crosses context boundaries
DIY multi-agent frameworks (Claude Agent SDK, CrewAI, LangGraph)
Pros
- Full control of brand voice, prompts, integrations, memory, oversight
- Reuse of existing Claude Code skill library (skills → agent tools)
- Rides on the LiteLLM proxy already on the roadmap
- Costs scale linearly with use, no per-seat markup
- IP and data stay in-house
- Same mental model the founder already uses daily with Claude Code
Cons
- Upfront engineering effort (weeks, not days)
- Ongoing maintenance burden (model changes, API changes, prompt drift)
- Oversight tooling is on you — approval queues, review dashboards, cost alerts
- Higher operational risk if not guardrailed properly
- Hallucination management is your problem, not the vendor's
8. Buy vs Build Analysis
What buying gets you
A working product today. No engineering work, no maintenance, no on-call. You trade control for speed: the vendor chooses the models, the prompt structure, the integration surface, and the update cadence. If Sintra adds a helper, you get it. If they raise prices or sunset a feature, you eat it.
For Objectuve specifically, buying breaks down in three places:
- Brand voice is extremely specific.
docs/brand/brand.mddefines a "coach who's also a friend" voice, anti-social-app philosophy, playful-but-not-childish tone. Vendors can't encode this without custom prompts, and custom prompts on vendor platforms are fragile. - The Claude Code skill library is a sunk-cost asset. ~40 skills, each containing Objectuve-specific constraints and patterns. A vendor can't read them. Building can.
- LiteLLM proxy is already happening for other reasons. The governance layer — cost tracking, model routing, fallbacks — is already being built for the consumer-facing AI Coach. Adding an AI Workforce on top is marginal effort, not a new project.
What building gets you
Every reusable piece Objectuve already has, wired together:
- Skill library (
.claude/skills/) → agent role definitions - LiteLLM proxy (
docs/product/completed/dedicated-ai-service-prd.md) → model routing + cost tracking - Sidekiq → scheduled agent runs
- Rails + GraphQL → artifact queue and approval dashboard
- Vue + Ionic → admin UI for oversight (exists; just add views)
- PostHog MCP, Sentry MCP, Mailtrap → tool access for agents
- Clerk → operator authentication on the approval dashboard
The marginal engineering effort is the orchestration layer, data model, and review UI — not a greenfield LLM integration.
The hybrid recommendation
Buy for velocity on low-stakes work. Build for brand-critical work.
Buy track: Pick Gumloop or Zapier Agents. Use for scheduled digests (weekly PostHog funnel summary, Sentry error triage, social inbox routing), platform posting (Buffer-style scheduled social), and lead capture handling. Ship within days. Budget: ~$37–$60/mo.
Build track: Start the AI Workforce service on top of LiteLLM. First employee: Content Drafter using
crafting-page-messaging+tightening-brand-voice. Ship in shadow mode (generates artifacts, human approves) within a few weeks. Expand to lifecycle email and SEO next.
The two tracks don't compete — they cover different risk profiles. The buy track handles operations where the downside of a bad output is "I look dumb on LinkedIn for an hour." The build track handles operations where the downside is "brand voice drifts and users feel the difference."
9. Build-It-Yourself Implementation Sketch
This is the "deep dive" portion. It's a sketch, not a spec — enough to confirm the approach is feasible and estimate cost, not enough to start coding.
Framework choice: Claude Agent SDK (committed)
Other candidates were CrewAI, LangGraph, and OpenAI Agents SDK. The recommendation is Claude Agent SDK, for four reasons specific to Objectuve:
- Skill library reuse. Claude Agent SDK's tool-use-first model maps directly onto the existing
.claude/skills/library. Each skill becomes either a tool the agent can invoke or a sub-agent it can delegate to. No translation layer. - Deepest MCP integration. Objectuve already uses MCP patterns in Claude Code (PostHog, Sentry, context7, firecrawl). The Agent SDK consumes the same MCP servers. Integration is literally a config change.
- Agents-invoke-agents. The "manager delegates to specialist employees" pattern is a first-class primitive, not a workaround. A
MarketingManagerAgentcan callContentDrafterAgentandSocialSchedulerAgentas tools. - Same mental model the founder already uses. Claude Code is already the daily driver. Building on the Agent SDK means the same mental model, skill authoring conventions, and debugging approach transfer directly. No new paradigm to learn.
CrewAI is a strong second choice if future contributors are more comfortable with Python than TypeScript/Ruby. LangGraph is the right choice if workflows become genuinely graph-structured with cycles — not yet the case.
Architecture
┌──────────────────────────┐
│ Operator Dashboard │
│ (Vue admin view) │
└───────────┬──────────────┘
│ GraphQL
┌───────────▼──────────────┐
│ Rails API │
│ - AiEmployee CRUD │
│ - AiRun status │
│ - AiArtifact approval │
└───────────┬──────────────┘
│
┌─────────────────────┼─────────────────────┐
│ │ │
┌─────────▼─────────┐ ┌─────────▼─────────┐ ┌────────▼──────────┐
│ Sidekiq │ │ Postgres │ │ Agent Runner │
│ (scheduled runs) │ │ (employees, │ │ (Claude Agent │
│ │ │ runs, │ │ SDK process) │
│ │ │ artifacts, │ │ │
│ │ │ memory) │ │ │
└─────────┬─────────┘ └───────────────────┘ └────────┬──────────┘
│ │
│ triggers runs │ LLM calls
└──────────────────────────────────────────┤
│
┌──────────▼──────────┐
│ LiteLLM Proxy │
│ (cost tracking, │
│ model routing, │
│ fallbacks) │
└──────────┬──────────┘
│
┌──────────▼──────────┐
│ Anthropic / OpenAI │
│ Gemini, Ollama │
└─────────────────────┘Data model sketch
# All inherit from PublicRecord, all soft-deleted with acts_as_paranoid.
class AiEmployee < PublicRecord
# Identity and role definition
# - name ("Cori, Content Drafter")
# - role_key ("content_drafter")
# - skill_refs (array, points at .claude/skills/ entries)
# - model_preference ("claude-sonnet-4-5" via LiteLLM)
# - schedule (cron expression, nullable for on-demand)
# - autonomy_level (shadow / semi_autonomous / autonomous)
# - active (boolean)
has_many :ai_runs
has_many :ai_employee_memories
end
class AiRun < PublicRecord
# One execution instance
# - ai_employee_id
# - triggered_by (schedule / manual / webhook / another_agent)
# - status (queued / running / succeeded / failed / awaiting_approval)
# - started_at, finished_at
# - token_count, cost_usd (pulled from LiteLLM)
# - error (if failed)
belongs_to :ai_employee
has_many :ai_artifacts
end
class AiArtifact < PublicRecord
# What the run produced
# - ai_run_id
# - kind (draft_post / email_template / code_patch / report / recommendation)
# - payload (jsonb — content, file path, PR link, whatever)
# - approval_status (pending / approved / rejected / auto_approved)
# - reviewed_by (user_id)
# - reviewed_at
belongs_to :ai_run
end
class AiEmployeeMemory < PublicRecord
# Long-term state per employee
# - ai_employee_id
# - key ("last_social_post_engagement_by_day")
# - value (jsonb)
# - updated_at
belongs_to :ai_employee
endSkill library reuse
Each file in .claude/skills/ becomes a reusable role definition. An AiEmployee record points at one or more skills via skill_refs, and the agent runner loads them as system-prompt fragments when the employee starts a run. tightening-brand-voice gets applied as a mandatory output filter on every employee whose output touches user-facing surfaces.
Integration points
- PostHog — read-only MCP tool, used by the analytics employee to pull funnel data
- Mailtrap — send-draft tool for the lifecycle email employee
- Sentry MCP — error triage tool for the customer support employee
- GitHub MCP — the content/SEO employee can open PRs against
marketing_landing/ordocs/ - Social platform APIs — deferred to Phase 2; Buffer or Typefully as the first target
- Clerk admin events — input signal for customer support (new signup → welcome sequence check)
Phased rollout
Phase 1 — Shadow mode. Every employee runs on schedule but produces artifacts only. All output sits in the approval queue. A human reviews, approves, and ships manually. Goal: calibrate trust, catch hallucinations, refine prompts. Measure: how often do artifacts ship unchanged vs require edits vs get rejected?
Phase 2 — Semi-autonomous. For roles where shadow mode showed >80% auto-approval rates, allow auto-publish within guardrails. Example: SEO meta-tag updates auto-merge to a draft branch but still require manual merge to master. Lifecycle emails get scheduled but not sent without approval. Social posts get queued in Buffer but not published.
Phase 3 — Scheduled autonomous. Roles with proven track records publish on schedule without review. Oversight shifts from per-artifact approval to anomaly detection: Sentry spans per run, budget alerts from LiteLLM, weekly brand-voice audit via tightening-brand-voice skill. A human still reviews the dashboard every morning but isn't a blocker on each run.
10. Cost Modeling
Rough ranges. Every number has a 50% margin of error; these are order-of-magnitude, not quotes.
Buy-only
| Item | Monthly cost |
|---|---|
| Sintra X ($97) OR Gumloop Pro (~$37) + Lindy Plus (~$50) | $37–$97 |
| Zapier Team (for integrations) | $70–$100 |
| Total | ~$100–$200/mo |
Velocity: days to first value. Ceiling: execution gap and credit caps. Best for low-stakes automation.
Build-only
| Item | Monthly cost |
|---|---|
| LiteLLM proxy (Cloud Run, already budgeted elsewhere) | Marginal |
| LLM token costs (10 employees × 50 runs/mo × avg $0.05/run) | ~$25–$100 |
| Additional Cloud Run compute for agent runner | ~$15–$40 |
| Additional Postgres storage (tiny) | Marginal |
| Engineering time (amortized, one-time) | Weeks of focus |
| Ongoing maintenance | ~2 hrs/week |
| Monthly run cost (after build) | ~$40–$150/mo |
Velocity: weeks to first value. Ceiling: none except engineering bandwidth. Best for brand-critical work.
Hybrid (recommended)
| Item | Monthly cost |
|---|---|
| Gumloop Pro (low-stakes automations) | ~$37 |
| Build track run costs | ~$40–$150 |
| Total | ~$80–$190/mo |
Velocity: days for the buy track, weeks for the build track. Best of both.
11. Risks & Mitigations
| Risk | Likelihood | Mitigation |
|---|---|---|
| Hallucination. Agent writes plausible-but-wrong copy, code, or data. | High (15–20% on complex queries per 2026 research) | Shadow mode Phase 1, approval queue, automated brand-voice audit via tightening-brand-voice, diff review on all code/config changes |
| Brand-voice drift. Over time outputs feel "AI-generated" and stop matching Objectuve's coach-friend tone. | Medium | Mandatory tightening-brand-voice post-filter on every user-facing artifact. Monthly voice audit by founder. |
| Cost runaway. Scheduled runs explode token usage. | Medium | LiteLLM per-employee budget caps. Alert at 80% of monthly budget. Hard stop at 100%. |
| Silent failure. Agent thinks it shipped but the artifact never reached its destination. | Medium | Sentry spans per run, per-step instrumentation, Slack alert on any run where finished_at - started_at deviates >2σ from baseline |
| Over-reliance / skill atrophy. Founder stops reviewing outputs and misses subtle brand/strategy drift. | Medium | Keep humans in the loop on strategic decisions. No Phase 3 autonomous mode for roles that set direction (growth hypotheses, roadmap bets). |
| Vendor lock-in (on the buy track). Gumloop/Lindy workflows become business-critical and impossible to migrate. | Low | Keep buy-track workflows to low-stakes, replaceable jobs. Document every workflow so it can be rebuilt in ~a day. |
| Documented precedents. Medvi's customer service chatbot fabricated drug prices that the company honored, and hallucinated product lines. | — | Shadow mode is the only safe default for customer-facing roles. |
12. Recommendation
Do this, in this order:
1. Buy track — ship this week
Subscribe to Gumloop Pro (~$37/mo). Build three low-stakes automations:
- Weekly PostHog funnel digest posted to a private Slack channel or email
- Sentry error triage (pattern-match new errors against last 30 days, flag anomalies)
- Inbox routing (tag and triage incoming marketing/support emails)
Why Gumloop and not Sintra: real execution (not just drafts), visual debugging, no credit ceiling, meaningful integrations into the existing stack. Skip Sintra entirely — its persona layer isn't worth $97/mo when the underlying models are cheaper via LiteLLM and the brand voice wrapper is weaker than tightening-brand-voice.
2. Build track — start within the month
Pin the LiteLLM proxy extraction as a hard dependency. Once that's live, build the AI Workforce service on top using Claude Agent SDK.
First employee: Content Drafter. Role: generates blog post drafts, landing page copy revisions, and feature release narratives. Skills loaded: crafting-page-messaging, framing-release-stories, tightening-brand-voice. Ships in shadow mode only. Success metric: >60% of generated drafts ship with <10% edits after 4 weeks.
Second employee: Lifecycle Email Designer. Role: drafts welcome series, re-engagement, and milestone emails. Skills loaded: designing-lifecycle-messages, tightening-brand-voice. Shadow mode only. Output lands in Mailtrap as unsent drafts.
Third employee: SEO Auditor. Role: weekly audit of marketing landing meta tags and Schema.org markup; opens PRs against marketing_landing/ with recommendations. Skills loaded: inspecting-search-coverage, adding-structured-signals. PR-based workflow = shadow mode by default.
3. Decide after 6 weeks
At 6 weeks, review:
- Buy-track: which Gumloop automations are you actually using? Kill the rest.
- Build-track: what's the auto-approval rate on shadow-mode artifacts? If >80% on any employee, graduate it to Phase 2. If <40%, rework the prompt/skill combination.
- Cost: total AI workforce spend vs baseline founder hours saved. Honest accounting.
If the build track delivers, add employees 4–6 (Customer Support, Analytics Reporter, Onboarding Auditor). If it doesn't, the buy track is still paying for itself at $37/mo and the build-track code isn't wasted — it becomes the consumer-facing AI Coach evolution path.
What to explicitly NOT do
- Don't buy Sintra. Persona wrappers over commodity LLMs aren't a moat, and the 250-credit ceiling punishes real use.
- Don't skip shadow mode. Hallucinations are documented and costly. The approval queue is non-negotiable for customer-facing surfaces.
- Don't build the full workforce at once. One employee at a time. Prove shadow-mode approval rates before adding the next role.
- Don't wire auto-publish before Phase 2 criteria are met. Brand drift happens gradually and is hard to reverse.
13. Next Steps
Framed as experiments, not dates. Run them in sequence.
Week 1 experiment — Gumloop spike
- Sign up for Gumloop Pro trial
- Build the PostHog weekly digest automation
- Measure: does the first digest ship something actually useful, or does it need 3+ iterations?
- Decision gate: useful → keep, add Sentry triage. Not useful → cancel, revisit Zapier Agents.
Month 1 experiment — Content Drafter in shadow mode
- Pin LiteLLM proxy as prerequisite. Status tracked in
docs/product/completed/dedicated-ai-service-prd.md. - Build the minimum viable
AiEmployee+AiRun+AiArtifactdata model - Wire one employee: Content Drafter
- Produce 5 draft artifacts, review each, measure approval rate and edit distance
- Decision gate: approval rate >60% → add Lifecycle Email Designer. <60% → refine prompts and skills, rerun.
Quarter 1 experiment — Three-employee workforce, Phase 2 candidate
- Content Drafter, Lifecycle Email Designer, SEO Auditor all running in shadow mode
- Build the approval dashboard (Vue admin view + Rails GraphQL)
- Graduate the first employee that hits >80% auto-approval to Phase 2 (semi-autonomous)
- Publish a follow-up doc:
docs/operations/ai-workforce-retrospective.mdwith real numbers — approval rates, cost per output, hours saved
Appendix — External sources cited
- Sintra AI Review 2026 — Social Rails
- Sintra AI Review 2026: The Execution Gap — Zaturn AI
- Sintra AI Pricing 2026 — Social Rails
- Sintra Reviews on Trustpilot
- 7 Best Sintra AI Alternatives — Gumloop
- Top 7 Sintra AI Alternatives 2026 — Lindy
- Lindy vs Gumloop Comparison — Inkeep
- Relevance AI Pricing — Lindy
- Agent SDK Overview — Claude API Docs
- Definitive Guide to Agentic Frameworks 2026 — Softmax Data
- Best Multi-Agent Frameworks 2026 — GuruSup
- AI Hallucinations Research — Kanerika
- AI Hallucinations Top User Concerns — AI Daily
Appendix — Internal references
docs/features/ai-coach.md— the existing consumer-facing AI surface (distinct from AI Employees)docs/product/completed/dedicated-ai-service-prd.md— the LiteLLM proxy that the build track rides ondocs/operations/vendors.md— current vendor inventory; this doc is linked from the "Under Evaluation" section theredocs/brand/brand.md— brand voice constraints that shape vendor fitrails_api/app/services/ai/coach_service.rb— the current LLM integration pattern for reference.claude/skills/— the specialized skill library that serves as the build-track's job-description layer