← Back to home
Feature 02

When AI Starts Connecting an End-to-End Finance Workflow

OpenAI Strategic Finance Days 7–12: From Cross-Surface Synchronization and Forecast Adjustments to Organizational Memory

Days 7–12 move beyond showcasing point tools. They connect refresh, analysis, review, and delivery within the same scenario, while turning models, content, and approved responses into reusable assets. The conditions for forming a finance workflow chain are now in place, but these capabilities do not yet constitute a unified finance operating system.

Days 1–6 broke finance work into executable units. Days 7–12 ask the next question: once those units begin running continuously, how do numbers, narratives, outputs, and organizational memory stay aligned across tools?

Introduction

During the first six days of #12daysofChatCodexStratfin, OpenAI’s Strategic Finance team showed how AI can enter specific segments of finance work: collecting operating signals, organizing cross-system data, generating candidate options, building analytical pages, performing reconciliations, and preparing month-end close materials. AI did not directly replace budget approval, accounting judgment, or management sign-off. Instead, it first transformed steps that had long depended on Excel, email, manual transfers, and individual experience into repeatable units of work.

The focus shifted in Days 7–12.

The question was no longer simply whether AI could complete a task. It became: after the P&L is refreshed, do Sheets, dashboards, and slides use the same version? After a monthly forecast is disaggregated to daily granularity, can it still reconcile to the original plan? When model recommendations incorporate account signals, are they double-counting information already embedded in the baseline? When a set of materials is converted into audio, are the numbers and qualifications still preserved? After AI generates a multi-tab long-range planning model, do the formulas, scenarios, and cross-tab logic actually close? Can an approved investor response be reused correctly and compliantly in the future?

These six days fall into two groups:

  • Days 7–9 address consistency in operation. The process must maintain the same operating facts across multiple outputs, time grains, and types of business signals.
  • Days 10–12 address reusable organizational assets. Outputs no longer stop at a one-time report; they become content, models, and knowledge that can be transformed, refreshed, retrieved, and accumulated.

If the keyword for Days 1–6 was “work units,” Days 7–12 show how those work units begin extending upstream and downstream, creating the conditions for them to connect into a workflow chain. A point tool can make one person faster. But unless adjacent steps gradually share common inputs, versions, evidence, and review states, Finance will merely migrate from six sets of manually maintained files to six conflicting AI tools.

The analysis below continues to examine each day through the same set of questions: What is the business bottleneck? What inputs and context are required? What does AI do? What responsibilities remain with people? What reusable asset is created? And what is the recommended starting scope?

Group One: Days 7–9, Keeping Numbers, Time, and Judgment Aligned

Day 7 | P&L refresh: Let multiple finance outputs share the same refresh

Business problem

During the close and forecast update, Finance rarely maintains only one spreadsheet. When the plan changes in Anaplan, the team must also update Sheets, the CFO dashboard, variance charts, commentary, and executive slides. The same set of numbers is copied into multiple outputs, each with its own refresh sequence, formulas, and owner.

The risk in the traditional process is not merely that it is slow, but that “everything looks updated, yet the outputs are not actually on the same version.” New Databricks actuals may be paired with an old Anaplan forecast; Sheets may be refreshed while slides still reference the previous version; the numbers may match while the period, scenario, currency, or narrative does not. In the end, manual checks usually see only the outputs and cannot confirm which source snapshot each page actually used.

Day 7 frames the problem as Source → Refresh → Analyze → Explain → Audit. It attempts to organize the updates and checks across multiple outputs into a single controlled run, instead of having analysts maintain each file separately.

Inputs and context

At minimum, this work unit requires an approved Anaplan snapshot, a Databricks extract, a shared period/scenario/as-of, a controlled P&L/metric view, output mappings, and a versionable set of review rules.

In a saved comment reply, the author explicitly stated that Skills were used. This matters: review steps no longer exist only as habits in an individual analyst’s workflow and can potentially be encoded as a reusable playbook. But a Skill becomes a control asset only when it has a version, owner, tests, tolerances, and execution records. If it is merely an opaque set of instructions, it can still produce incorrect results consistently.

AI’s role

According to the author, Codex handles the refresh, analysis, explanation, and audit handoff. Reconstructed as a production workflow, a more robust implementation would have the system first check sources, schemas, and control totals, then generate a controlled P&L/metric view. Sheets, dashboards, and slides would all refresh from that view instead of copying from one another.

AI would then compare variances, identify the main drivers, and draft commentary using approved business context or meeting notes. Finally, versioned rules would check the numbers, period, and narrative, placing failed items into an exception queue.

The public GIFs show Ready for review, P&L AUDIT CHECK, and AUDIT READY. These states are valuable because they distinguish automated prechecks from handoff to human review. Here, AUDIT READY should be understood as “review preparation is complete” or “the materials are ready for human review,” not “AI has completed an audit,” and certainly not “ready to publish.”

Human role

The Finance owner still selects the official source version, confirms metric definitions, investigates exceptions, and determines whether the narrative reflects business reality. The reviewer needs to inspect the evidence and exceptions produced by the Skill, not just a row of green checkmarks.

People are also responsible for testing the review Skill itself. The team should inject errors into a test workbook: an incorrect period, a stale scenario, a sign error, a formula overwritten with a value, inconsistencies across outputs, and a missing footnote. If the Skill cannot detect known errors, it cannot serve as the gate for AUDIT READY.

Reusable asset

The most valuable asset from Day 7 is a cross-surface reporting pipeline:

  • A common source snapshot and controlled metric layer;
  • Output mappings for Sheets, dashboards, and slides;
  • A versioned review Skill;
  • Consistency checks for numbers, periods, units, scenarios, and narratives;
  • An exception queue, reviewer decision, and publish version;
  • The source, output, and Skill hash for every run.

This allows Finance to stop updating files one by one and instead run a reporting refresh with defined inputs, states, and evidence.

Do not begin with the full CFO pack. Select four line items—Revenue, Gross Margin, R&D, and Operating Profit—and connect one Anaplan snapshot, one Databricks extract, one Google Sheets workbook, one dashboard page, and two slides.

The first version must prove three things: all outputs come from the same snapshot; the Skill detects injected known errors; and any change to a number or commentary revokes the prior review status. Once those three conditions are met, expand to more metrics and pages.

Scope of public evidence: Two GIFs directly show the refresh, CFO reporting pack, P&L audit checklist, and audit handoff. The actual connectors, Skill content, control totals, cross-output tie-outs, and approval records were not made public.

Day 7 P&L refresh and audit handoff

Caption: The key to Day 7 is not automatically refreshing more pages, but ensuring every page shares the same refresh—and treating AUDIT READY in the UI as the start of human review.

Day 8 | Ads forecast: Disaggregate the monthly plan into a daily operating model

Business problem

Many finance models plan by month, while operating actions occur daily. Ads revenue is affected by weekdays, holidays, seasonality, regions, and user trends. A monthly forecast can support management planning, but it has difficulty answering how traffic, impressions, and revenue should change on a particular day.

Teams previously maintained both monthly and weekly forecasts. Whenever assumptions or business logic changed, they had to disaggregate, aggregate, and reconcile the figures again. The real challenge is not dividing a monthly number by the number of days. It is ensuring that daily, weekly, and monthly views use the same assumptions and reconcile in both directions across month boundaries, partial weeks, holidays, and rounding.

Inputs and context

The workflow requires an approved monthly plan, an actuals snapshot, the model/scenario/as-of, a weekday profile, a holiday calendar, country/product mappings, and an explicit rounding policy.

The calendar should not be treated as an ordinary parameter. Whether a holiday uplift overrides the weekday effect, how a week spanning two months is allocated, and how incomplete actuals are displayed will all change the results. Every weighting and precedence rule needs to be versioned; otherwise, the same month may produce a different daily plan depending on the refresh date.

AI’s role

Using approved weekday, holiday, and seasonality rules, the system disaggregates the monthly forecast into weekly and daily operating plans and generates daily, weekly, and monthly views along with a budget-versus-actual comparison. The disaggregation formulas, bidirectional reconciliation, and treatment of rounding residuals must be deterministic; the model cannot be allowed to reason freely on each run.

The public GIF shows a Daily Ads Model, Ad-Eligible DAU, Billable Impressions, Ad Revenue & ARR, multiple cases, time-grain switching, and actuals comparisons. This indicates that the case does more than output a disaggregation table; it builds a forecasting app that can compare versions, scenarios, and time grains. AI can help build the application, compare scenarios, identify drivers, and explain exceptions.

The production version should also use a deterministic checking layer to perform bidirectional tie-outs after every disaggregation: daily aggregates to weekly, weekly aggregates to monthly, with subtotals also checked by country, product, and plan type. Rounding residuals must not be silently discarded; they should be explicitly allocated and recorded.

Human role

Finance and business owners define the weekday, holiday, and seasonality logic, decide which cases may enter the official forecast, and explain actual deviations. The system can calculate multiple scenarios, but people still approve the official version.

People must also decide how the results enter downstream systems. A commenter asked directly whether adjustments are written to an ERP/FP&A cube in real time from the console or staged for review first. The author did not answer publicly. For most companies, the first version should generate only a staged export. Writing back to the official planning model or FP&A cube should wait until approval, least privilege, duplicate-write protection, writeback reconciliation, and rollback capabilities are all in place.

Reusable asset

Day 8 creates a time-grain forecast engine:

  • A monthly plan and actuals snapshot;
  • A weekday/holiday calendar and disaggregation rules;
  • Bidirectional tie-outs across daily, weekly, and monthly views;
  • A versioned scenario store;
  • BvA, overrides, exceptions, and approval history;
  • A staged export and rollback pointer.

It brings the monthly finance plan into day-to-day operations, instead of maintaining a separate operational sheet disconnected from the official forecast.

Select one metric, one country, two cases, and a three-month plan. Write out weekday and holiday weights explicitly, and inject month-boundary, missing-holiday, stale-actuals, and rounding errors.

The first version is sufficient if it can reliably generate daily and weekly tables, fully roll them back up to the monthly total, and retain every exception in a review queue. After approval, save an immutable version and output it to a staging table without directly rewriting the official cube.

Scope of public evidence: A 13-frame GIF directly shows a synthetic forecasting app, case comparisons, Daily/Monthly switching, and BvA. The weekly view, disaggregation algorithm, tie-outs, code, tests, and writeback were not made public.

Daily and monthly views of the Day 8 Ads forecast

Caption: Turning a monthly forecast into a daily operating model is not primarily about finer granularity. The core requirement is that different time grains still reconcile to the same approved version.

Day 9 | Account-level forecast adjustment: Put models and Finance judgment into the same review chain

Business problem

Data Science models can provide an account-level baseline, but the latest business changes often appear first in Gong, Slack, RevOps records, spend signals, or product launch information. If Finance ignores these signals, the forecast will respond slowly. If it simply layers every signal onto the baseline, it may double-count information already incorporated into the model.

The real problem in Day 9, therefore, is not “letting AI change the forecast for Finance.” It is allowing the model baseline, prior forecast, business evidence, and human judgment to be compared, explained, and saved in the same interface.

Inputs and context

Inputs include the DS model/version, prior forecast, point-in-time account signals, account master, event time and ingestion time, source entitlements, and rules for determining whether each type of signal has already been incorporated into the baseline.

Two foundational controls are required. The first is point-in-time integrity: information arriving after the forecast cut-off cannot be backfilled into the recommendation as it stood at that time. The second is cross-source deduplication: the same customer event may appear in Gong, Slack, and RevOps and cannot be counted as three separate incremental events.

AI’s role

The system first performs point-in-time filtering, entity matching, cross-source deduplication, and baseline-inclusion checks to avoid future-information leakage and double counting. AI then identifies drivers from unstructured calls, messages, and account events, organizing the DS baseline, prior forecast, recent trend, and multi-source business signals into an account-level review. It generates a recommendation with source evidence, confidence, and proposed impact.

The public GIF shows Human review required, Apply recommendation, version saving, and XLSX/Sheets exports. In a more mature workflow, Apply should not directly rewrite the official forecast; it should send the recommendation into an accept, modify, or reject queue. Every reviewer decision should record a reason, after which the versioning system creates a new version and deterministic calculations update the quarter, FY, and prior-forecast bridge.

AI’s most valuable role is reducing the time Finance spends finding account context, interpreting unstructured signals, and drafting adjustment recommendations. It cannot simultaneously serve as signal extractor, forecast owner, and final approver.

Human role

The Finance reviewer determines whether a signal truly changes the forecast, confirms whether the model has already absorbed similar information, and documents the reason for an override. RevOps or the account owner corrects account mappings and business context.

People must also challenge numbers such as Confidence 86%. Without a calibration set and historical performance, confidence is merely the model’s internal expression and cannot directly become an approval threshold. High-value adjustments should require stronger evidence and a higher approval level even when confidence is high.

Reusable asset

Day 9 creates a forecast recommendation and override system:

  • A versioned DS baseline and prior forecast;
  • Account events recorded by type and with point-in-time snapshots;
  • Baseline-inclusion and cross-source deduplication records;
  • Recommendations with citations;
  • Accept/modify/reject decisions and override reasons;
  • An immutable forecast version, export hash, and downstream reconciliation.

This turns forecast adjustments previously scattered across models, chat logs, and analyst judgment into a reviewable decision history.

Use ten synthetic accounts and one quarter. Fix the DS baseline and prior forecast, then create Gong, Slack, and RevOps events containing duplicates and future-dated records.

The first version must pass point-in-time filtering, entity matching, baseline-inclusion, and cross-source deduplication tests. Only recommendations with original citations should enter human review. Export to a staging workbook should be allowed only after the version has been written to a controlled version store.

Scope of public evidence: The public GIF shows adjustments, evidence, recommendations, human review, versions, and exports. Source records, schemas, deduplication, point-in-time tests, persistence, and downstream writeback were not made public.

Day 9 account-level forecast adjustment

Caption: Day 9 does not use business signals to replace the statistical model. It puts the baseline, evidence, human adjustments, and version history into the same review chain.

Group Two: Days 10–12, Turning Outputs, Models, and Responses into Organizational Assets

Day 10 | Document/deck → podcast: Convert analysis into a more accessible executive briefing

Business problem

Finance and Data Science teams may spend weeks completing an analysis, only to post a deck in Slack. Managers do not have time to read every page, and important conclusions, qualifications, and action items can easily get buried in the attachment.

Day 10 addresses insight delivery. It does not produce a new forecast. Instead, it converts an approved document or presentation into a podcast-style briefing so managers can absorb the key information during a commute or between meetings.

Day 10 extends the consistency of the same facts across different delivery formats. The same approved materials can produce a new format, but they cannot produce a second set of facts: the source deck is the source of truth, while the script and audio are derived versions. Numbers, periods, and qualifications cannot be lost during the conversion.

Inputs and context

The primary input should be an approved source file, not an arbitrary working draft. The system also needs the audience, length, tone, voice, pacing, an approved TTS project, and the numbers, dates, units, ranking methodology, scope exclusions, and footnotes that must be preserved.

An executive script is more than a summary. It needs to know which caveats cannot be removed during compression, which chart titles cannot be separated from their denominators, and when conflicts between speaker notes and tables must be escalated to a person.

AI’s role

The original post says the Skill identifies the core narrative and key conclusions, then converts them into podcast-style audio. For formal executive materials, the process should be split further: AI extracts the narrative, numbers, and caveats and drafts a script with slide/page citations; deterministic checks verify numbers, units, dates, and scope qualifications; the content owner approves the final script; and TTS is used only to generate the audio. After generation, the system compares the transcript with the approved script and sends pronunciation or omission issues to review.

The public video directly shows a finished podcast of approximately 111 seconds, whose content is broadly consistent with the central facts and qualifications in OpenAI Q1 2026 Signals. But it does not show the source deck, Prompt, Skill, script, TTS parameters, or human approval. The existence of a finished product whose facts are broadly consistent does not mean the generation process is auditable.

Human role

The content owner approves the script and confirms that key numbers and caveats have not been weakened in pursuit of fluency. The reviewer must also check the pronunciation of names, acronyms, currencies, percentages, and technical terms.

For a confidential deck, Security/IT must confirm which API project receives the file, whether retention is allowed, and whether the output audience matches the source ACL. The distribution scope for the audio cannot be broader than that of the original file simply because “it is only a summary.”

Reusable asset

Day 10 creates a governed executive briefing pipeline:

  • An approved source and version hash;
  • A claim/number/caveat inventory;
  • A script with citations;
  • Script approval and a change log;
  • A fixed TTS model, voice, instructions, and speed;
  • An audio-transcript diff, pronunciation review, and distribution ACL.

What is truly reusable is not a particular voice, but the complete lineage from source to script to audio.

Use a five-page synthetic Finance deck containing ten numbers, one ranking table, two footnotes, one scope exclusion, one speaker note that conflicts with the body, and several difficult-to-pronounce names. Generate a 90-second briefing.

It should pass only if every number, unit, period, and qualification is preserved; conflicts are escalated; the script is approved by a person; the audio transcript matches the script; and the output includes an AI-generated disclosure.

Scope of public evidence: The original post directly states that the team built a reusable Codex Skill to convert a document or presentation into a podcast, with adjustable voice, tone, pacing, and audience. The video, audio track, subtitles, and transcript prove that the final podcast exists, and its core facts can be cross-checked against the official OpenAI Q1 2026 Signals page. A script with page-level citations, numerical checks, human approval, transcript diffing, and pronunciation review are part of the production reconstruction proposed in this article; the public materials do not show these steps.

Final Day 10 podcast output

Caption: Day 10 demonstrates a new format for management to receive information. Its production value depends on whether a complete evidence chain is preserved across the source, script, and audio.

Day 11 | AI-native long-range planning model (LRP): Let AI generate the model—and make the model subject to finance controls

Business problem

At a minimum, a long-range planning model needs to keep products, scenarios, periods, revenue, COGS, OPEX, working capital, and the P&L linked. If the model is further used for a full budget, financing, or board planning, it will generally need to extend to cash flow and the balance sheet and complete a three-statement tie-out.

AI can generate a multi-tab workbook in a short time, much faster than building it cell by cell. But “the spreadsheet is filled in” is not the standard for a completed model. Broken links, hardcodes, sign errors, unit mismatches, scenarios that do not propagate, and broken cross-tab logic can all hide inside a workbook that appears complete.

Inputs and context

Day 11 starts with a detailed Prompt and a blank workbook. A production version also needs versioned requirements, a data dictionary, product/scenario configuration, a source snapshot, model architecture, a formula policy, color conventions, named-range conventions, and acceptance tests.

Before modifying an existing model, AI should first propose a change plan: which sheets will be created, which ranges will be overwritten, and which existing formulas must be preserved. Otherwise, a large-scale generation run can easily damage the existing model without explaining what changed.

AI’s role

AI builds the workbook structure, helper columns, named ranges, hidden support tabs, formulas, a flattened Model Dataset, and an Outputs Pivot. The public video shows the process from a blank Google Sheet to P&L Modeling and an Outputs Pivot, indicating that AI can already generate a complex model structure.

An independent checking layer must then run formula, link, sign, unit, period, scenario-propagation, Source Key uniqueness, dataset-to-source reconciliation, and pivot-to-dataset reconciliation tests. If the company extends the model to cash flow and the balance sheet, it should add a three-statement tie-out. The agent that generates the model should not declare it passed solely on the basis of its own summary.

For live data, AI may read only from a controlled staging/import tab and must save the source system, query/version, extracted-at, as-of, row count, and control total. If a refresh fails, publication must be blocked; the system cannot silently continue using stale actuals.

Human role

The FP&A owner defines drivers, scenarios, and business relationships, reviews key formulas, and compares the AI-generated version with an independent baseline. Accounting or a Finance reviewer checks working capital and cross-tab logic and, if the model is expanded to three statements, owns the corresponding tie-outs. The Data owner manages source refreshes and lineage.

People must also decide when the model can move from prototype to formal use. Generating a sanitized example in a few hours does not mean it is ready for budgeting or board planning. Speed matters only when the model’s structure, formulas, tests, versions, and review can all be replayed.

Reusable asset

Day 11 creates a model factory with controls:

  • Versioned requirements and a detailed build Prompt;
  • Workbook architecture and formula conventions;
  • A source import contract and lineage;
  • Automated formula, reconciliation, and scenario tests;
  • Traceability from source tabs → Model Dataset → Outputs Pivot;
  • Reviewer sign-off, workbook/source hashes, and controlled export.

This turns AI from a one-off spreadsheet writer into a model-building tool constrained by tests.

Use a synthetic model with three products, three scenarios, and eight quarters, and deliberately add a broken link, hardcode, duplicate Source Key, stale actuals, and a unit mismatch. If the pilot explicitly extends to a complete three-statement model, also add an error such as an “unbalanced balance sheet.”

The pilot passes only when automated tests detect every injected error, the dataset and pivot reconcile, key outputs match an independent baseline, and the reviewer, workbook, and source hashes are saved. Full three-statement closure applies only to an expanded version that explicitly includes cash flow and the balance sheet.

Scope of public evidence: A 31-second video directly shows the upper portion of the detailed Prompt, a blank start, the multi-tab model, P&L Modeling, and the Outputs Pivot. The complete Prompt, workbook, formulas, tests, three-statement tie-out, live refresh, and human review were not made public.

Day 11 multi-tab LRP workbook

Caption: Day 11 shows that AI can rapidly generate a complex workbook structure. Finance teams still need independent tests to prove that it is a model, not merely a set of tables that looks complete.

Day 12 | IR Agent: Turn approved responses into organizational memory

Business problem

Investor diligence is not ordinary knowledge Q&A. The same question may permit the use of different materials at different financing stages, for different investors, and at different points in time. Numbers are updated, narratives are superseded, and some information may be disclosed only to specific recipients.

The traditional process depends on a small number of people remembering the history: how the team answered last time, which numbers were approved, what content was later replaced, and which materials can be shared with whom. As request volumes grow, the team can duplicate work and may also reuse a response that is outdated or unsuitable for the current recipient.

The core of Day 12 is not to build a “chatbot that can answer investor questions.” It is to organize requests, sources, drafts, approvals, delivery, and final responses into continuously updated institutional memory—that is, organizational memory.

Inputs and context

The system needs an authorized investor request, the investor/deal/stage, recipient entitlements, a point-in-time approved source set, and Legal, Comms, and executive approval policies. Information governance should be split into two dimensions: base access levels of Public, Internal, Confidential, and Restricted; and separate disclosure restriction tags such as MNPI, investor-specific, deal-stage-limited, and legal-review-required. MNPI is not an access level parallel to Public or Confidential.

Approved responses cannot simply be appended to a vector store. Every record needs effective-from/to, supersedes/superseded-by, applicable recipients, prohibited reuse conditions, source passage, claim, number, period, as-of, and approver. Retrieval must first filter on base permissions, then deterministically filter by disclosure tags, investor, financing stage, and effective date, and only then perform semantic ranking.

AI’s role

According to the author, the IR GPT drafts investor diligence responses from source materials and writes the approved final response back to a curated knowledge base. It was later expanded to analyze requests, create investor materials, and support other investor inquiries. To use this method in a formal IR process, a deterministic system must first filter by entitlement, disclosure tags, investor, financing stage, and effective date. AI can then retrieve from the allowed source set, generate a draft with passage-level citations, and check numbers, dates, names, as-of dates, conflicts, and staleness.

The final response can be sent only after the required Legal, Comms, or executive approval. Only the final response that was actually approved and sent may be written back to the curated knowledge base together with its approval record.

The public screenshot shows the IR Agent, Google Drive, the investor-reporting Skill, a ChatGPT/Slack channel, and partial role instructions. At minimum, this demonstrates the existence of the corresponding agent configuration. The original post supports drafting responses from source materials and writing approved responses back to the knowledge base. But passage-level citations, entitlement filters, effective dating, delivery hashes, and a supersession schema are enterprise designs proposed in this article; the screenshot does not directly show operational evidence for these mechanisms.

Human role

The IR owner is accountable for every external response and checks citations, numbers, wording, and applicability. Legal, Comms, and executives handle MNPI, selective disclosure, material commitments, and financing language in accordance with policy.

People are also responsible for maintaining the validity period of the knowledge. One approval does not mean permanent approval. When a new financial period, financing stage, or company narrative appears, old answers need to expire or be explicitly marked as superseded by a new version. Otherwise, the feedback loop can amplify a historical error into “organizational memory.”

The author cited 2,000+ diligence requests, 10,000+ estimated hours saved, over $180 billion of capital raised, and no financial advisor engaged. The public materials do not include request logs, time records, or an attribution methodology. These figures describe the author’s account of the three-person team’s overall scale of work across multiple financing cycles. They do not represent the IR Agent’s independent incremental contribution to the amount raised and cannot be treated directly as agent ROI.

Reusable asset

Day 12 creates an approved-response knowledge system:

  • Requests, investors, stages, and entitlements;
  • A point-in-time source set and passage citations;
  • A claim/number/date/as-of map;
  • A draft-to-final diff and approval chain;
  • Controlled delivery and an artifact hash;
  • Effective-dated approved responses;
  • Supersession, reuse scope, and activity logs.

It means organizational memory is no longer merely “making old answers easier to find.” The system can establish when an old answer applied, to whom, on the basis of which source, and with whose approval.

Do not directly replicate a financing scenario. First build a synthetic diligence room with four access levels—Public, Internal, Confidential, and Restricted—and add tags such as MNPI, investor-specific applicability, and financing-stage restrictions to selected materials. Then add two financing stages, one superseded answer, two sets of conflicting as-of figures, and 50 requests containing duplicates and adversarial inputs.

Acceptance requirements include zero MNPI leakage, 100% claim-level citations, exclusion of all expired responses, escalation of every conflict, repeatable entitlement filtering, full logging of review decisions, and admission of only approved final responses into the knowledge base.

Scope of public evidence: A static screenshot directly shows the agent, Google Drive, Skill, files, channel, and partial instructions. The actual requests, retrieval, citations, approvals, sends, knowledge writeback, complete Prompt/Skill, and underlying KPI records were not made public.

Day 12 IR Agent configuration

Caption: The most important part of Day 12 is not “remembering more answers,” but ensuring that every answer enters organizational memory with its sources, permissions, validity period, and approval record.

Phase Two Conclusion: Finance Needs Shared Conditions for the Workflow Chain

Days 7–12 appear to cover reporting refreshes, Ads forecasting, account adjustments, podcasts, LRP, and investor relations. Remove the application names, and they all address four common requirements.

First, every output must trace back to the same controlled facts

Day 7 requires Sheets, dashboards, and slides to come from the same P&L refresh; Day 8 requires daily, weekly, and monthly forecasts to reconcile in both directions; Day 9 requires an account recommendation to trace back to the DS baseline, prior forecast, and original business evidence; Day 10 requires audio to trace back to the approved script and source deck; Day 11 requires the pivot to trace back to the dataset, formulas, and source rows; and Day 12 requires an investor answer to trace back to a currently effective source passage that is permitted for disclosure.

The infrastructure for finance AI, therefore, is not just a connector. It also needs source snapshots, metric definitions, versions, hashes, citations, and lineage. Without them, AI can generate six kinds of output faster but cannot prove that they are saying the same thing.

Second, state matters more than “done”

Traditional automation often has only running, success, and failed states. Finance work needs more granular states: source approved, refreshed, exception, review ready, reviewed, approved, staged, published, and superseded.

AUDIT READY does not mean audited. Apply recommendation does not mean written back. Generated audio does not mean its script was approved. A populated workbook does not mean the model passed. Retrieving an old response does not mean it is still permissible to use.

A mature workflow chain must make clear which step has been completed, who owns the next step, and which changes invalidate the current state.

Third, human judgment needs to leave a structured record

The review of exceptions in Day 7, scenario approval in Day 8, accept/modify/reject decisions in Day 9, script approval in Day 10, model sign-off in Day 11, and external-response approval in Day 12 should not exist only as a “looks good” message in a chat.

At minimum, Finance should save the decision, reason, reviewer, timestamp, and corresponding artifact version for each human judgment. This lets the organization distinguish between AI’s recommendations and decisions for which people accept accountability, and later review which recommendations were accepted, modified, or rejected.

Fourth, true reuse comes from rules, tests, and memory

Day 7 encapsulates review rules in a Skill; Day 8 turns calendar logic into a disaggregation engine; Day 9 turns overrides into version history; Day 10 turns content conversion into a briefing pipeline; Day 11 turns a detailed Prompt and tests into a model factory; and Day 12 turns approved responses into effective-dated knowledge.

These assets can all operate across cycles. But reuse also amplifies errors: a flawed Skill repeatedly lets issues pass, an incorrect holiday rule continuously distorts the daily plan, and an incorrect response is recalled repeatedly from the knowledge base. Reusable assets must also have an owner, version, tests, known limitations, and retirement/supersession mechanisms.

Practical Implications for Finance Leadership

Looking only at the screens, Days 7–12 can easily be interpreted as “the OpenAI finance team built six more AI applications.” The more important change is that Finance’s objects of management are beginning to shift from files and individual tasks to a set of work units capable of continuous operation and the potential connections among them.

Finance Leadership needs to establish four shared foundations:

  1. Controlled data and definitions: Common source snapshots, as-of dates, scenarios, a metric registry, and mappings;
  2. Workflow states: Exception, review, approval, staging, publish, and supersession;
  3. Evidence and testing: Control totals, formula checks, citations, lineage, known-error tests, and output hashes;
  4. Accountability and permissions: Owners, reviewers, approvers, recipient entitlements, and rollback.

This also changes project priorities. Teams should not begin by pursuing “one agent that covers all of FP&A.” A better sequence is to select one high-frequency workflow chain, fix its inputs and approval boundaries, establish an exception queue and independent tests, run it for several consecutive cycles, and then connect adjacent steps.

For example, a team could begin with the four P&L metrics in Day 7, first connecting Sheets, a dashboard, and two slides. Once stable, it could add the executive audio briefing from Day 10. Or it could begin with ten accounts from Day 9, first recording recommendations and human overrides. Once stable, it could feed the approved forecast version into the daily operating model from Day 8. The sequence of connections should be determined by business dependencies, not by which demo is more visually compelling.

How Far Days 1–12 Have Progressed

From AI4FIN’s analytical perspective, Days 1–12 can be arranged as a progressively expanding path:

Operating signals and finance inputs
→ Executable work units
→ Consistency across tools and time grains
→ Exception, review, and approval states
→ Reusable models, content, and knowledge

Days 1–3 expand the sensing layer, allowing Finance to see marketing, sales, and workforce changes earlier. Days 4–6 turn analytical pages, reconciliations, and the close deck into stateful work units. Days 7–9 keep multiple outputs, time grains, and multi-source signals aligned. Days 10–12 begin turning executive outputs, long-range planning models, and approved responses into organizational assets.

At this point, the twelve days have established a set of finance capabilities, but not yet closed the loop on a final operating model. The public cases have separately demonstrated the target forms and working methods for P&L refreshes, daily forecasts, account adjustments, podcasts, LRP, and an IR knowledge base. But they have not yet answered several larger questions: How are these capabilities combined? How are they reused across teams? Who owns routing, handoffs, schedules, and exceptions? And how do the outputs of local agents enter a unified Finance management system?

Nor do these cases prove that AI can already operate a finance function independently. The public materials rely heavily on synthetic, sanitized, or illustrative data, while complete Prompts, Skills, schemas, tests, permissions, and production logs are rarely disclosed. What they prove is something more specific:

These cases show that AI’s application boundary is no longer limited to answering questions alongside Excel. It is beginning to work with deterministic systems across the inputs, refreshes, exceptions, reviews, publication, and memory of finance work. Financial calculations, tie-outs, permission filtering, and state transitions remain constrained by formulas, code, and rules systems, while people’s responsibilities shift toward defining standards, challenging models, handling exceptions, approving scenarios, and accepting final accountability.

But “multiple reusable capabilities have emerged” is not the same as “a Finance-managed workflow chain has formed.” From directly modifying a single dashboard, to analytical engines that can repeatedly answer new questions, to a Treasury workflow and a cross-organizational Agent Network, the subsequent Bonus Days would continue to add the layer of capability composition, specialized agents, orchestration, and governance.

Days 7–12 are therefore better understood as the endpoint of a second phase, not the conclusion of the entire series. The Bonus Days would continue to answer one final question:

Once models, content, knowledge, and work units can all be reused, how does Finance organize them into a system of work that can operate, hand off, review, and assign accountability across teams?


Primary Sources

Source Scope

This article is based on saved Day 7–12 LinkedIn posts, images/GIFs/videos from the main posts, samples of public comments, and official OpenAI materials available as of July 27, 2026. The public materials were used to reconstruct the working methods; they do not indicate that OpenAI released complete data, code, Prompts, Skills, tests, or production configurations. Figures disclosed by the authors for time savings, usage scale, and capital raised remain attributed claims in the absence of public underlying records and are not presented as independently verified results.