A note from the founder  ·  November 2025 – present

What a year of building agentic AI in Indian finance operations actually taught me.

Six design partners. A go-to-market that failed for a reason worth knowing. The pivot it forced. And one customer who decided, on their own, to pay me ten times more. Here is all of it, including the parts that didn't work.

6 design partners 2 ICPs tested ~86,000 lines of code Price fit found
01 — The shift

From systems of record to systems of activity

The belief I started with, and still hold: the value is moving from where data is stored to where work is decided.

Every piece of finance software built in the last thirty years is a system of records. Tally records transactions. SAP stores ledger entries. These systems are passive — they wait to be queried. They don't know the close is three days behind, that a vendor was paid twice, or that this quarter's TDS is being computed under the wrong section.

The human is the active intelligence. The accountant opens the report, notices the anomaly, chases the exception, closes the month. The software is a warehouse. The human is the worker.

records sit still …the human does everything the system works surfaced for judgment SYSTEM OF RECORDS SYSTEM OF ACTIVITIES
Same ledger, same data. The difference is who carries the work — and what reaches the human.
The closest analogy is CRM. In 2000, Salesforce was a contact database. By 2015 it logged every call, prompted follow-ups, flagged deals going cold. The salesperson's job changed from managing a spreadsheet to managing a queue of prioritised actions.

Finance operations is ten years behind sales in making that shift. From my product charter, June 2026

In the CRM transition, the activity layer captured the value — not the record layer. In Indian finance the record layer is already spoken for: Tally holds 90%+ of SMB books. The activity layer is not yet built. That is the bet I have been making with my time.

02 — The operating principle

Signal over noise

The product is not the automation. The product is the exception.

Out of three hundred reconciled transactions, the value sits in the thirteen that got flagged — each with a plain-language reason and a recommended resolution. The other 287 are a log, not a view.

300 transactions reconciled  ·  13 need a human
Almost everyone builds the screen for the 287. The work — and the money — is in the 13.

This wasn't a UX preference I argued for. It was something a real client's statutory books proved to me.

1,564Tally ledger rows ingested — journal, purchase and purchase-GST registers
56 / 56In-scope Form 26Q deductions matched. Nothing left for manual matching
47Auto-resolved from previously captured human decisions, out of 63 learned rules
₹1,21,397Undeclared TDS liability found under §194H — in books that had already passed the client's own compliance process

The matching engine confirmed everything was fine. The checker found the money. Three vendors sat in Brokerage & Commission ledgers with no corresponding Form 26Q entry — exposure carrying 1%/month interest under §201(1A) and potential 30% disallowance under §40(a)(ia).

A reconciliation tool finds mismatches between two declared records. What creates real exposure is the record that isn't there — and finding an absence means reasoning forward from the expense head, to the section that should have applied, to the threshold, to the silence where a filing should be. That reframing changed what I thought I was building.

03 — Research, chapter one

Finding the problem

Who is in enough pain to change how they close their books — and who can actually decide to pay for that change?

Before writing a line of agentic workflow code, that was the question I kept coming back to. The answer moved once, and moving it is the most useful thing I learned all year.

I mapped Indian SMBs by how they actually staff accounting at each revenue band — not how they should, but what I saw on the ground.

1 visit / month 1 visit / week on site, most days daily · 1 dedicated daily · in‑house < ₹1 Cr ₹1–5 Cr ₹5 Cr+ ₹5–15 Cr ₹15 Cr+ Small kirana —daily essentials Larger store —e.g. cloth retail Modern retail Modern retailwith own brand D2C brand —strong digital presence ANNUAL REVENUE HOW THE BOOKS ACTUALLY GET DONE JR ACCOUNTANT ON SITE the decision point paying for a dedicated person — not yet their own hire
Below the highlighted band the pain isn't acute enough to justify change. Above it, the business has already solved it by hiring in-house.
The staffing staircase — what I saw on the ground
Annual revenueThe businessWho does the data entryHow often
< ₹1 CrSmall kirana, daily essentialsThe CA sends a junior accountant to collect bills and receiptsOnce a month
₹1–5 CrLarger store — e.g. cloth retailA junior accountant sits in the shop and enters bills into Tally on siteWeekly
₹5 Cr+Modern retail1–2 outsourced staff dedicated to the businessOn site, most days
₹5–15 CrModern retail with its own brandOne fully dedicated accountant — still outsourced, not yet in-houseDaily
₹15 Cr+D2C brand with strong digital presenceThe business hires its own full-time in-house accountantDaily, in-house

Somewhere between ₹5 and ₹15 crore, a business is paying for a dedicated resource but hasn't yet decided to own the function. Real recurring spend on bookkeeping, real dependency on someone who isn't their employee, and no owned system of record for how that work gets done. That's where I aimed.

To reach them I went through the top 300 CA firms — not as the end customer but as the channel, chosen for the AI-implementation muscle that separates them from the median firm. That choice wasn't about reach. Three things made a firm a far better place to stand than any single business.

01
Pain visible at scale

One firm carries 200–400 clients. Standing behind a single firm meant watching the same pain repeat across hundreds of businesses at once — rather than discovering it one company at a time, one sales call at a time.

02
Capex to build the brain

A firm that size can fund building an AI capability for its own team, and has someone inside who can run it. A single ₹5–15 Cr business has neither the budget nor the person, and won't grow one for this.

03
A direct line to outcomes

The firm sees what actually happened to the end client — what got filed, what got disallowed, what the auditor flagged, what the notice said. That feedback loop is the only thing that makes a system smarter month over month.

Taken together: my bet was that a firm capable of running a pilot moves faster than the businesses themselves, and teaches you more per conversation.

What I actually tested

Not a survey — a real go-to-market motion over roughly six months. Outbound to 300+ large firms, pitching at the AI Summit at Bharat Mandapam, and POCs at VSN&Co, HNA (Hiregange & Associates), Guru & Jana, Kuhoo Finance and Taxpert.

04 — The POC portfolio

Five proofs of concept

Not pitch decks. Five firms handed over real books and real people, and I built against each of them.

They are listed and dated by when each engagement began. Each POC targeted a different finance workflow, identified through a customer journey discovery workshop run with the firm — mapping how work actually moved through their team, where it stalled and who it stalled on, before anything got built. Together — after months working alongside their teams, not weeks of interviews — they are the reason I can describe this market from the inside rather than from a survey.

THE PIVOT CONTINUES IN 06 MetaMorphoSysNewGen VSN&CoGuru & Jana HNA & Co AI CFO DASHBOARDAI TREASURY TDS RECONSALES-TO-CASH GST COMPLIANCE …and here the journey changed NOV 2025 — AUG 2026
Five POCs inside the CA-firm thesis — five different workflows, five sets of real books, and the same ending each time. What lies past the dashed line is section 06.
Nov 2025
MetaMorphoSys
AI CFO Dashboard

The first build, and where the thesis met a screen. A CFO cockpit with a structured query contract: every answer has to arrive carrying the evidence rows it was derived from and an explicit statement of what it could not determine. The reasoning layer, not the frontend, chooses which KPIs and charts answer the question.

6,408 lines of code4 required response fields

What it taught me: the layer from ledgers to MIS to dashboard is broken at every joint — each asset sitting with a different person, each maintaining their own version. A dashboard built on top of that cannot answer with atomic visibility; it shows numbers nobody can trace back to a transaction. The problem was never the chart.

Jan–Apr 2026
NewGen
AI Treasury — fixed deposits

The full lifecycle of bank fixed deposits for a multi-entity group across four jurisdictions — booking, approving, accruing, posting. Maker, Checker and CFO roles, with approval tier derived from the deal's own economics rather than picked from a dropdown. Control enforced at the database: six Postgres roles whose grants make a forbidden write impossible even if the application layer has a bug.

26,571 lines of code52 tables, 6 role-bound engines2 full UAT rounds

What it taught me: controls and decisions belong to the user; agents are for task execution. Enforce that split at the lowest possible layer — and never let a user declare their own authority.

Mar–Apr 2026
VSN&Co
TDS reconciliation

Form 26Q declarations reconciled against a live company's Tally books, section by section, across a full financial year. A five-pass matching cascade, then a separate compliance checker that reasons forward from the expense head to the section that should have applied — and flags the filing that was never made.

1,564 ledger rows ingested56 / 56 in-scope matched₹1,21,397 undeclared TDS found

What it taught me: the exception is the product. Matching confirmed everything was fine; the checker found the money.

May–Jun 2026
Guru & Jana
Sales-to-cash & GSTR-2B

The agentic build. An articled-clerk agent that reads Tally, reconciles GSTR-2B against the purchase register, chases missing vendor confirmations over WhatsApp, drafts the corrective journal entries and queues them for partner sign-off. Every tool declares a risk tier; the ones that write to the ledger don't execute at all until a partner types CONFIRM.

13,751 lines of code4,620 lines of specification23 tools, 3 approval tiers

What it taught me: the buyer is the partner. Replace the grunt hours, never the judgment — and brand every vendor message as the firm, never as us.

Jun 2026
HNA & Co
GST compliance lifecycle

Lekha AI for GST compliance — the whole monthly and quarterly lifecycle in one loop: GSTR-1, GSTR-2B reconciled against the purchase register, GSTR-3B, and challan creation. The 2B reconciliation is the painful part and mandatory monthly for ITC claims, but the value is in never leaving the loop. The build's answer to hallucinated compliance rules: a curated registry of versioned statutory rules, each with a citation, an effective-date range and a named owner. Where no rule fits, the agent must escalate rather than reason from training data.

10+ rule categories encoded100% of answers carry a citation

What it taught me: for a CA firm to be genuinely empowered, the end client has to be the ultimate user of the system — an answer that stops at the firm's desk never reaches the person whose books it concerns. And the human-in-the-loop thesis held: citations made every answer checkable, and a person still made the call.

Interest was high at every one of them. All five ended on the same sentence.

The turn
From custodian to owner
The pivot
CA FIRM PARTNER holds the books · carries the risk owns none of the upside the turn CFO, QUICK-COMMERCE SELLER owns the books · owns the decision feels the deduction

Same thesis. A different person on the other side of the table. The specifics are in 07.

05 — Channel learnings

Why Indian CA firms aren't ready for agentic systems yet

This isn't a hedge. It's the thesis I'd defend in a room, and I'd genuinely like to be argued with on it.

HIGH INTEREST MetaMorphoSysNewGenVSN&Co Guru & JanaHNA & Co “We will not share live client data with a third-party AI system.” THE IDENTICAL BLOCKER — EVERY FIRM, NO EXCEPTION five doors opened…
Five distinct workflows, five sets of real books, one repeated refusal. That single, identical no is the actual finding.
01
Letting go of control is a rational fear

Agentic adoption is still genuinely early here, and the reluctance is not ignorance. These firms have watched good sellers end up at the mercy of Amazon and Flipkart — dependent on a platform that later set the terms. Handing an agent the keys to a client's books looks, from where they sit, like the same trade.

02
The maturing cycles are unfunded capex

An agent is not finished when it is built. Consistent, predictable outcomes come from running it over and over against one client's real context — and those cycles are where the AI brain actually matures. That phase needs AI engineers on payroll for months. It is capex, not a subscription, and nobody in this chain is funding it.

03
In-house now looks genuinely possible

With Claude Code on the desk, a capable firm can build its own workflows — and every week that works raises their confidence that doing it in-house is real rather than theoretical. That growing confidence was the most effective competitor I met, and it isn't a product I can outbuild.

THE THREE BLOCKERS, IN THREE WORDS 010203 WON’T CAN’T NEEDN’T hand a third party the keys to a client’s books fund the months of maturing cycles an agent needs buy it at all — Claude Code is already on their desk this is the one I can’t outbuild
The first two are objections you can work on. The third is a competitor getting better every week, and it isn’t a product — it’s their own growing confidence.
One adjacent pricing learning

India is a genuinely different market to sell AI-native services into. You cannot price at pre-AI, legacy-bookkeeping rates while you are the one running and owning the service end to end with AI. A subsidised entry price is a real option — but it is a strategic bet that needs funding behind it to burn.

06 — Research, chapter two

Choosing the vertical

Move the ICP from the custodian to the entity that owns its own books — then work out which kind of business that should be.

The refusal in chapter one told me what to stop doing. Before committing to any segment I ran the field through four filters, each one drawn from something the CA-firm chapter had already cost me to learn. Anything that failed a filter was set aside, however large the market looked.

Every Indian business running Tally Owns its own books Pain is deadline-bound, not discretionary Enough volume that doing it by hand hurts Data arrives in a form a machine can read no custodian between the pain and the decision a statutory date, or a claim window with a clock on it not a handful of invoices a month portals and payout files, not a shoebox of paper D2C · e-commerce · quick-commerce sellers four filters, each one already paid for
Every filter here is a lesson from the previous chapter turned into a test. The first is the custodian problem. The second is why TDS converted and payment reconciliation didn't. The third and fourth are what separates a business worth automating from one where a person with a spreadsheet is genuinely cheaper.

D2C and e-commerce brands selling through quick commerce — Blinkit, Zepto, Instamart — and the marketplaces clear all four. They own their books outright, their deductions come with claim windows measured in working days, the line-item volume is punishing, and every platform already publishes the data through a portal. I kept CA firms as a referral and credibility layer, not as the channel.

07 — The thesis

Quick commerce, specifically

Sellers get paid net of deductions, and the deductions never arrive with the sale.

The pain, specifically

Sellers get paid net of deductions, and those deductions never land in the same accounting period as the sale. Six deduction types arrive from Blinkit, Zepto and Instamart across four different systems, on four different days — and each platform does it its own way. Left unmanaged, books lag reality by two to three weeks.

D0 D+7 D+14 D+21 D+30 D+45 Short supply RTV Expiry rejection Damage Ad spend Listing fees Sale ONE SALE · SIX DEDUCTIONS · FOUR SYSTEMS nothing lands in the same period Illustrative example of a platform payment cycle
The deduction never arrives with the sale. Reconciling it is the whole job — and the clock is running the entire time.

It's getting structurally worse, not staying flat. Blinkit moved to a 1P inventory model on 1 September 2025 — reconciliation now happens at PO ↔ GRN ↔ invoice ↔ debit-note level rather than order level, and every short supply legally requires a GST credit note. Instamart is heading the same way. Zepto's RTV process is its own trap: it arrives as a tax invoice on the vendor requiring a purchase voucher, but the original sale never reverses — quietly inflating both sales and purchases and breaking P&L reconciliation if nobody catches it.

And there's a hard clock. Miss the claim window and the deduction is simply unrecoverable.

CLAIM WINDOWS — WORKING DAYS TO SCALE MARKETPLACES NykaaFlipkartAmazon 7 days14 days60 days QUICK COMMERCE BlinkitZeptoInstamart window to confirmwindow to confirmwindow to confirm the deduction is booked to the period the incident is recorded — not the period of the sale 01020 30405060
Miss the window and the deduction is unrecoverable. Nykaa gives you seven working days. The three quick-commerce tracks stay unfilled on purpose: I went looking and their seller-side claim windows aren’t published outside the seller portal, so I won’t draw a bar to scale from a number I can’t source. What is documented is worse anyway — on quick commerce the deduction is booked to the settlement period in which the incident is recorded, not the period of the sale, so the row you are chasing isn’t even in the month you are closing.

Which segment, and why

I ranked segments by where the deduction pain is structurally worst, not biggest in absolute size.

1
Health & nutra
FSSAI's 30%/45-day shelf-life rule turns expiry-driven RTV into a mandated, recurring monthly debit-note stream. Highest rupee value per line item.
~37%all-in take rate
(agency estimate)
2
Beauty & personal care
The top 75 brands are roughly 75% of category ad spend, so the addressable list is short and knowable.
up to 10%of QC sales lost to
ad deductions (Redseer)
3
Packaged food & beverages
Highest line-item volume of any category, and the thinnest margins in the set — a 2% deduction error hurts here more than anywhere.
37 hrsa week reconciling, at
one ₹4 Cr brand
—
Apparel · parked
Returns are high but mostly third-party. A strong visibility problem, a weaker posting story. Later phase.
25–40%return rate
Each segment is ranked on its own signature number, in its own unit — a take rate, an ad-deduction share, hours a week, a return rate. They are not comparable to each other, and ranking them meant weighing what each one costs the seller rather than putting them on one axis.

The evidence, graded honestly

₹1,000 TRIAL MONTH ₹10,000 PER MONTH, ONGOING scope expanded at the same price they chose to spend more REVEALED PREFERENCE — ONE LIVE PAYING CUSTOMER Tally posting running end to end · reconciliation now running beyond credit notes
The strongest signal available at this stage isn't a survey response. It's a customer whose spend went up on their own initiative.

The pricing learning — the one I'd lead with next time

Pricing is the friction in AI adoption. Not capability. The buyer needs a value proposition they can compute, and if they can't compute it they wait.

I did not read this. I paid for it — across six design partners and more than fifty businesses, over a year of conversations that ended politely and went nowhere.

6design partners, five workflows, real books
50+businesses in the pricing conversation
1price that finally held — and the buyer raised it, not me
WHY AN AI PRICE STALLS — AND THE ONE QUADRANT THAT DOESN’T every box ends with the sentence you actually hear Software substitution The salary swap Act of faith Outcome pricing They price you against the licence they already pay. You only win on cheaper. Priced against a person they would stop needing. They never stop needing them. Nothing to compare, nothing to calculate. No number they can defend upward. We predict your AR and your cash deductions — the accounting is real time. “Tally already does most of this.” “So who exactly do we let go?” “Let’s revisit next quarter.” ₹1,000 → ₹10,000, unprompted THEIR NUMBER: A TALLY LICENCE THEIR NUMBER: A SALARY THEY KEEP WHERE EVERY AI PITCH STARTS THE ONLY QUADRANT THAT GETS PAID CANNOT COMPUTE THE BENEFIT CAN COMPUTE THE BENEFIT IS THERE A NUMBER THEY ALREADY PAY? YES NO three dead ends — and the exit is one square to the right
The vertical axis asks one concrete question: is there already a number on this buyer’s books that yours will be compared with — the Tally licence, the outsourced-bookkeeping retainer, the accountant’s salary? If there is, that number becomes your ceiling, whatever you are actually worth. Having none looks like a disadvantage, and it is, right up until the buyer can compute the benefit — at which point it becomes the advantage, because nothing caps you. The salary swap is the quadrant founders mistake for the safe one: it sounds concrete and it is the least credible thing on this chart, because nobody in that room believes the headcount is going anywhere. You are asking them to pay a real invoice against a saving they will never book. The three sentences are the archetypes I heard, not transcripts; the number in the fourth box is the one thing here that actually happened.
Pricing basisWhat the buyer has to believeVerdict
Per seatthe SaaS default
That more logins mean more value. In finance the opposite is true — the best outcome is fewer people touching the ledger.
Contradicts
Per transactionusage-metered
That volume equals value. It taxes the good months and prices the exceptions — the only rows that matter — the same as the 287 that don't.
Misprices
Flat subscriptionthe licence model
That your category is worth a line item. You are now compared with software they already own, on a number they already know.
Caps
Capital releasedoutcome-linked, with a cost-to-serve floor
Only their own arithmetic. Deductions recovered inside the claim window, and receivables that stop being unconfirmed. You did the conversion for them.
Pays
Read the middle column as the ask. The first three require the buyer to convert your unit into their value, and a finance team that is already sceptical about AI will not do that conversion for you. The fourth arrives already converted.

The mechanism, in one line

Deductions are not a reporting problem. They are a cash problem. An unreconciled deduction is a receivable you cannot confirm, and a receivable you cannot confirm is working capital you cannot deploy. Manage the deduction stream, AR becomes predictable, and predictable AR is capital released.

THE ARITHMETIC A CFO CAN DO IN THEIR HEAD ₹15 Cr₹4.1 L12 days ₹50 L annual revenueof revenue a dayof lag removed WORKING CAPITAL RELEASED ÷ 365 × books run two to three weeks behind; twelve days is the conservative end DRAWN TO THE SAME SCALE What it costsWhat it frees ₹1.2 L a year, at the price the live account pays today ₹50 L, once roughly forty times the fee — and that is the argument, not the demo
This is a thesis, modelled from the reconciliation lag established earlier in this section, not an audited outcome at a customer. I am showing the working so it can be argued with: ₹15 Cr of revenue is ₹4.1 lakh a day, and twelve days of lag removed is ₹49.3 lakh of capital that stops sitting in unconfirmed receivables. The number I would defend is the method, not the fifty.

What fifty conversations actually taught me

L1
If the buyer has to convert your unit into their value, they won't. They will say they'll think about it, and they will mean it.
L2
A price with no anchor is an act of faith — and finance teams are professionally disqualified from buying on faith.
L3
“Cheaper than a person” prices you against a headcount cut nobody actually makes. The saving never reaches the P&L, so neither does your invoice. Price the capital you release, not the labour you imagine displacing.
L4
The strongest pricing signal isn't a yes. It's a customer who raises their own spend without being asked.
The shape of the solution this implies

A config-driven, per-client agentic pipeline — not a generic rules engine. SKU and ledger mapping learned from each client's own historical vouchers rather than a one-size-fits-all chart of accounts. Every credit note drafted and reviewed before it posts; nothing writes to Tally without a human check. Failed postings become exception records with a suggested fix, never a silent retry.

08 — The hard part

Architecture of trust

Everyone is shipping an agent. The question is whose agent a finance team will let near the ledger.

The answer I arrived at is not a feature list. It is a claim about what is actually scarce.

A finance team can absorb a fixed number of interruptions in a day. That budget — not compute, not model accuracy — is the real constraint on any agent near a ledger.

Every component I built exists either to protect that budget, or to earn the right to spend one unit of it. The claim I would defend in a room

What earns the right to interrupt a human

Most of the work is deciding what never reaches a person. The rest is justifying the few times you do.

ONE CLIENT · MAY 2026 · GSTR-2B RECONCILIATION 120 invoices in 108 matched 12 exceptions purchase register + GSTR-2B deterministic match — no LLM a rule ID, a reason, one decision 90% of the file never costs anyone a moment of attention narrow too hard = a missed claim every band removed is attention given back
Counts are the committed fixture state for the one client the pipeline runs against — 120 purchase invoices, 108 clean matches, 12 exceptions across eight categories. The vertical note is the part that keeps this honest: narrowing is not free. Filter too hard and the thing you suppressed was a claim inside its window, which is why the matching path is deterministic rather than a model deciding what looks unimportant.

Two ways to keep the budget cheap

The 10,000-row rule. Would this agent still work if the file had ten thousand rows? If the LLM sits inside the iteration loop, no — cost and latency scale linearly with volume, which is an architecture mistake rather than a scale problem. The LLM takes the supervisor seat; tools do the iterating. That's why the reconciliation engine has no LLM in its matching path at all.

Rules the model is forbidden to invent. The dominant failure of LLMs in compliance is confident invention of a plausible rule. I answered it structurally: a curated registry of versioned compliance rules, each carrying a statutory citation, an effective-date range, a named owner and a review date. Every conclusion carries the rule IDs it applied. When no rule fits, the agent must flag the line as a gap and escalate. It may not reason from training data. Because rules are versioned by effective date, a historical period gets reconciled under the rules that applied then.

What one unit of attention actually buys

Here is a real exception from that run, at full size. Everything I believe about this product is visible in one card.

ONE EXCEPTION, ANNOTATED Auto Wheels India Pvt Ltd AWI-INV-2026-441 · 15 May 2026 · HSN 8703 INELIGIBLE ₹3,59,520 Motor vehicle — Toyota Innova, 7-seater. Input tax credit is not available. Reverse ₹3,59,520. gst-itc-sec-17-5-a-motor-vehicle · v2026.05.1 Sec 17(5)(a), CGST Act 2017 · effective 01 Jul 2017 owner CA Pradeep Guru · reviewed 15 Apr 2026 EVIDENCE Tally voucher PUR/2526/00467 — passenger vehicle, seating ≤ 13, so Sec 17(5)(a) applies. Type CONFIRM to post to Tally CONFIRM Reject Approve & post JE Partner only. The clerk persona cannot complete this. why, in a sentence the rule, versioned and owned by a person the voucher it came from no third option — nothing posts itself and not everyone may press it the whole philosophy fits on one card — that is the argument
Real row, real rule, real voucher reference — this is ex-005 from the committed reconciliation fixture, rendered as the product renders it. Note what the card refuses to do. It does not show a confidence score alone, it does not offer a “post anyway”, and the rule it cites carries a version, an owner and a review date so a partner can disagree with the rule rather than with the machine.

The rule behind the rules: autonomy is earned against reversibility

Three tests decide whether a feature ships — can it explain itself, can you stop it, does it ask before anything irreversible. But tests are pass/fail, and what I actually needed while building was a way to decide how much autonomy a given action had earned. This is that decision, drawn.

AUTONOMY IS EARNED AGAINST REVERSIBILITY EASY TO UNDO IRREVERSIBLE FULLYAUTONOMOUS SUGGESTSONLY the ceiling — nothing sits above it THE DEAD ZONE autonomous and irreversible — where agent products get banned fetch and normalise the files deterministic matching draft the credit note write to client memory send vendor comms post to Tally descope autonomy, not intelligence — as a placement rule, not a slogan
The diagonal is the whole policy: the further right an action sits, the less autonomy it may hold. Memory writes are the one people miss. They look harmless and sit mid-chart in amber, but a bad single action is visible and correctable while a bad memory rule silently corrupts every future run — so they are gated too. Most agent products let memory writes happen invisibly.

How the gate is enforced

A gated tool's handler doesn't execute and then ask forgiveness. A separate function builds the approval request without performing the action, so there is no code path where the side effect happens before a human sees it.

ExplainControl Confirm what & whypause · overrideexplicit sign-off HUMANINTENT AGENTACTS NOTHING REACHES THE LEDGER WITHOUT PASSING ALL THREE the gate a human holds
Explain, control, confirm. Three pass/fail tests, enforced in code across products built months apart.
If a feature fails any of these tests, the correct response is not to remove the AI. It is to add visibility until it passes.

Descope autonomy, not intelligence. Core product thesis
09 — How it fits together

Lekha, as an operating system

Every piece of finance software built in the last thirty years is a system of records. It waits to be queried. The human is the active intelligence in that arrangement.

Tally records your transactions. An ERP holds your asset register. None of them know your close is three days behind, or that a vendor payment went out twice, or that a quarter's TDS is being computed under the wrong section. They store what happened. Somebody has to open the report before anything is noticed.

The consequence of the last section is that this can be inverted. If attention is the scarce resource, the software should be the thing that works continuously and the human should be the thing that gets interrupted rarely — which is the opposite of how finance software has always been arranged.

THE SAME TUESDAY, TWICE 2 AM9 AM1 PM6 PM SYSTEM OF RECORDS download the files,open Excel, VLOOKUP scroll 120 rows lookingfor what broke still reconciling.CFO asks again nothing happens SYSTEM OF ACTIVITIES ran the recon andbuilt the queue 12 exceptions reviewed,the hard one escalated done before lunch the work moved to 2 AM — and it stopped being a person's work
The rhetorical point is that the top row is busy and the bottom row is empty. The emptiness is the product. A system of activities doesn't wait to be queried — it works in the background, surfaces what needs judgement, explains what it found, and stops before anything consequential.

The builds aren't separate products. They're the layers of one, discovered in the wrong order — modules first, platform last, because that's the order the design partners arrived in.

SOURCE SYSTEMS → ORBIT IN Lekha AI AGENTIC OPERATING SYSTEM TALLY/ZOHO TAXPORTAL PLATFORMPORTALS POS/SALES BANK INVOICES PAYROLL GST/TDSFILINGS CLOSED & DECISION-READY TRIAL BALANCEFINANCIALS MIS/REPORTINGCASH & PROFIT Your whole stack orbits Lekha. Decisions fall out closed.

Tally is desktop software; its only programmatic door is an XML gateway on localhost:9000, and a browser on an HTTPS page can't call it. That one constraint explains the shape of the whole market — why most tools shuttle CSVs, and why the cloud suites have to ask the SMB to migrate. I shipped a signed Windows connector, then documented ten reasons my own architecture was wrong, and rebuilt it to dial outward over TLS so automation runs without a human sitting at a browser.

Three separate products independently needed it. That convergence is what told me it's infrastructure rather than a feature.

10 — Next

What I want to do next

One workflow, taken all the way — the opposite of the last eight months.

The ICP is settled, the price fit is found, and the first customer is paying. The right move now is depth, not breadth: take quick-commerce deduction reconciliation from one live account to a repeatable product, in the segment where the pain is structurally worst and getting worse.

That means closing the loop end to end — platform deductions ingested, reconciled against the seller's own books, credit notes drafted and posted into Tally under human approval, exceptions surfaced in an activity feed with the claim-window clock visible. Nothing in that path needs an invention. Every piece is either shipped or specified. It needs hands, and a narrow enough focus to use them in one direction.

Alongside it, I'm continuing structured seller interviews across health/nutra and beauty to see how far the pricing holds across the segment.

If any of this is your world, I'd like to talk

If you run finance at a D2C or quick-commerce brand and the deductions above sound like your month — I'd genuinely like to hear how you handle them today, whether or not you ever become a customer.

If you're building in this space — reconciliation, book close, or finance agents for the Indian market — I have a year of expensive negative results and one hard-won ICP. Most of it is more useful to you before you spend the money than after.

If you're a larger company adding AI to an existing finance product — a compliance platform, an accounting suite, an order-to-cash or ERP player — what's here is a working Tally write-back layer, a validated deduction taxonomy for quick commerce, an agent safety architecture that survives a CA's scrutiny, and a map of which segments have budget and which don't.

And if you're just a founder who's curious how a year like this actually goes — that conversation is free, and I've been on the other side of it enough times to want to pay it back.

Appendix

Going deeper

Six product dossiers, each committed to its own repository. Every number on this page traces back to committed source.

P0
Tally Connector & Write-Back
tally-connector-installer

The shared substrate. The market barrier, my own v1 self-critique, the v2 outbound architecture, the Tier 1 write-back spec, read-side validation on 4,758 vouchers.

P1
TDS / GST Reconciliation
aitdsrecon

The engine, the five-pass cascade, the omission checker, and the full measured run against real client statutory filings.

P2
NexGen Treasury
newgen-v2

Derived approval tiering, six role-bound database engines, append-only audit, and the full process record.

P3
Lekha as an Operating System
AI-CFO-dashboard-v2

The charter, the three design tests, the 10,000-row rule, the cockpit query contract, and the CFO-facing narrative deck.

P4
Lekha Harness Agent
lekha-demo-client

The gated agent loop, the compliance registry, hash-chained audit, 334 traceability tags, and a trace-level evaluation framework.

P5
Credit Note Pipeline
nakad-tally

Seven steps, two gates, tokenized approval, real role enforcement — and the accounting question we refused to guess.