Set up your client project: brand voice guide, claim evidence, and the asset calendar
Teams build one Juma Project per client and add context over time. Every flow the team runs for that client pulls from the same project, so a caption run in month three starts from the voice profile, the approved claims and the posting patterns that months one and two established.
What to add
Brand Voice Guide
The brand's social voice in its own words: how posts open, how they close, the constructions that recur, the emoji posture, and the words to avoid. Add it and the voice profile step starts from the team's own reference instead of rebuilding it from scratch each month.
Claim Evidence Library
Impact reports, product pages, press releases and any published figure the brand stands behind, each with its date. This is what caption claims get mapped against, and it is what turns "we think this is fine" into a sourced line.
Asset Calendar
What is going out this month and what each asset actually shows. The more specific the description, the more specific the caption, and the less the copy defaults to generic product language.
Previous Caption Sets
Past months saved as reference. Recurring hooks, hashtags that keep appearing and formats that have been used recently all surface, so the new month varies rather than repeating itself.
Guide Juma with project info
Add a short description to each knowledge item in the project's info field so Juma knows what each file contains and when to use it. For example:
- Brand Voice Guide: "Our social voice. Check every caption against this before proposing it."
- Claim Evidence Library: "Every published figure we stand behind, with dates. Map caption claims to this and flag anything that does not map."
- Previous Caption Sets: "The last three months of captions. Avoid repeating hooks and hashtags that appear here."
Write this month's captions in one chat
Frequently Asked Questions
What does the social media caption playbook include?
The playbook includes a voice profile built from the brand's live site and public posts, one caption per asset per platform, the visible-before-cut text and character count for each, a pass or fail on the hook, tiered hashtags, functional alt text, and an approval table naming what to confirm per asset.
It arrives as a branded PDF plus an editable CSV with one row per asset per platform. The CSV is the source of truth: every count in it is recomputed rather than carried over, and the PDF is reconciled against it before delivery, so the two cannot disagree halfway through a scheduling session. The posting order for the month comes with a line of reasoning per slot rather than an arbitrary sequence.
How is this different from a generic social media caption generator, or from writing the posts themselves?
A generic caption generator writes from the brand name and a topic. This Flow starts from assets that already exist and writes the copy that goes with them, checked against three things a generator does not check: whether the hook survives the platform's truncation point, whether each claim maps to a sourced figure, and whether the alt text is functional.
That starting point is also what separates this from the Flows that build social content from scratch. Create a social media calendar, Create Instagram content and Write X threads all begin with a topic or a content pillar and invent the post. This one begins with the photograph already sitting in the folder. If the team needs to decide what to post, those Flows come first; if the assets are shot and the captions are what is missing, this is the one to run.
The truncation check is the one that changes the copy most. Character limits are widely known, but the number that decides whether someone taps "more" is the cut-off, not the cap: Instagram's feed hides text after roughly 125 characters regardless of how long the caption runs. Writing to the cap and writing to the cut produce different first lines.
How much time does this save compared to writing captions manually?
A month of five assets across five platforms is 25 captions. Written by hand with voice checks, character counting and hashtag research, that is most of a day. This Flow returns the full set with counts, alt text, claim flags and a scheduling sheet in minutes, and the social lead spends that day editing rather than drafting.
The saving compounds inside a Juma Project, because the voice profile and the claim evidence carry forward. Month two does not rebuild the voice, it checks against it. Strategy, taste and judgment stay human: which assets earn a post, which hook is actually funny, and whether a claim is worth making are decisions the team keeps. Juma augments the team, it does not replace it.
Does this work for brands that publish on only two or three platforms?
Yes. Name the platforms the brand actually publishes on and captions come back for those only. Each one is written to its own conventions rather than adapted from a master caption, so a two-platform run produces two genuinely different pieces of copy rather than one caption and a trim.
This matters more than platform count suggests. A LinkedIn caption and a TikTok caption for the same asset share a subject and almost nothing else: one leads with context and a takeaway, the other with a spoken first line built for on-screen pacing. Asking for every platform a brand does not use produces copy nobody publishes, and dilutes the copy for the platforms that matter.
Can the captions be written without a brand's social accounts being connected?
Yes. The voice profile is built from public sources, the brand's live site and its public posts, so no account connection or login is needed. Nothing is uploaded, and the Flow runs from the brand URL and the asset list alone.
Connecting nothing does have one limit worth naming: public posts show what a brand published, not which posts performed. If the team has performance data, adding it to the project as a knowledge item lets the posting order and hook choices reflect what has actually worked rather than what has simply been posted. Without it, the profile still captures voice accurately, since voice is visible in the copy itself. Human review on every output.