Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 6 additions & 6 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,7 +29,7 @@ Each message can sound harmless. Together, they create delivery work, approval r

The current prototype runs entirely in the browser and supports this flow:

1. Load a scope document and communication or work-record exports.
1. Load a scope document or initial order email and communication or work-record exports.
2. Normalize the local files into scope items and evidence records.
3. Run deterministic, configurable rules against the messages.
4. Show each potential scope-drift finding with its evidence, source and scope basis.
Expand Down Expand Up @@ -93,19 +93,19 @@ The current UI uses one local project at a time. **New local project** clears th

### 3. Add the source of truth

Use **Add a source** to upload a scope document and communication export. You can remove a source or correct its classification before running the analysis. The current supported formats are:
Use **Add a source** to upload a scope document or initial order email and a communication export. You can remove a source or correct its classification before running the analysis. If the agreed work is described in the first order email, select **Scope document** for that file in the source list and keep later correspondence as **Communication export**. The current supported formats are:

| Format | Typical use | Current support |
| --- | --- | --- |
| `.txt` | SOW, agreement, proposal or brief | Supported |
| `.md` | Markdown scope or delivery notes | Supported |
| `.md` | Markdown scope, delivery notes or RTF order export saved with an `.md` extension | Supported |
| `.eml` | One email message from Gmail, Outlook, Apple/iCloud, Proton, Yandex, Mail.ru or a custom domain | Supported |
| `.json` | Telegram, Facebook Messenger, Slack-style or other message export | Supported |
| `.csv` / `.xlsx` | CRM or ERP table export | Planned adapter |
| `.pdf` | PDF agreement or export | Planned adapter |
| `.docx` | Word agreement or export | Planned adapter |

File type is inferred from the filename and content. Names containing terms such as `sow`, `scope`, `contract`, `agreement`, `brief` or `proposal` are treated as scope documents. Names containing `slack`, `email`, `telegram`, `whatsapp`, `messenger`, `facebook`, `message`, `thread`, `chat`, `linear` or `jira` are treated as communication sources.
File type is inferred from the filename and content. Names containing terms such as `sow`, `scope`, `contract`, `agreement`, `brief`, `proposal`, `order`, `request`, `intake`, `kickoff` or `requirements` are treated as scope candidates. Names containing `slack`, `email`, `telegram`, `whatsapp`, `messenger`, `facebook`, `message`, `thread`, `chat`, `linear` or `jira` are treated as communication sources. You can always correct the classification in the source list.

For the most reliable result, name files clearly. For example:

Expand Down Expand Up @@ -166,7 +166,7 @@ This keeps the open-source project useful and auditable while protecting the int

Click **Run analysis** after both source types are loaded. ScopeGuard then:

- extracts bullet points from the scope source;
- extracts bullet points, clear prose deliverables or readable RTF order fields from the scope source;
- extracts messages from the communication source;
- applies the configured rules;
- calculates the current finding list, the share of messages with a scope basis, message count and preliminary exposure hours.
Expand Down Expand Up @@ -267,7 +267,7 @@ The intended operating habit is simple: review the queue before accepting new wo
- Direct Slack, Gmail and WhatsApp integrations are not implemented yet; export adapters are available and live pilot connectors are the next open-source slice. Facebook Messenger, additional mailbox providers, CRM and ERP integrations are not part of the open-source pilot.
- The JSON parser supports common Telegram, Facebook Messenger and Slack-style arrays and objects with a `messages` array. It looks for fields such as `text`, `message`, `body` or `content`, including Telegram's array-of-text-fragments format.
- WhatsApp text exports are treated as message sources when their filename includes `WhatsApp`, `chat` or a similar channel hint.
- Scope extraction currently focuses on Markdown-style headings and bullet or numbered list items.
- Scope extraction supports Markdown-style headings, bullet or numbered list items, clear prose in an initial order email and RTF order exports saved as `.md`.
- Rule matching is lexical and deterministic. The pilot now handles basic negation, sender role hints, included/excluded clauses, stable finding IDs and multiple rule matches, but it does not understand full contractual context or conversation history.
- “With scope basis” is the share of parsed messages for which at least one configured rule found a related scope clause. It is not a percentage of contractual compliance.
- Direct API integrations, shared workspaces and project switching remain outside this pilot slice.
Expand Down
4 changes: 3 additions & 1 deletion docs/implementation-notes.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@ The current slice is a responsive React/Vite local-first review workspace, a det

- No external UI library yet; the surface is small and the visual language is custom.
- Sources are normalized into `SourceDocument` records with a scope/messages/unknown kind and an explicit format.
- TXT/MD, EML and common Telegram, WhatsApp, Facebook Messenger and Slack JSON exports are parsed locally in `src/analysis.ts`.
- TXT/MD, RTF-in-MD, EML and common Telegram, WhatsApp, Facebook Messenger and Slack JSON exports are parsed locally in `src/analysis.ts`.
- The analyzer runs deterministic rules for new deliverables, acceptance criteria, extra revisions and unpriced commitments.
- Rule patterns, rule signal strength, severity and hour ranges are configured in `src/rules.yaml` and validated at build time.
- “Run analysis” now recomputes findings, the percentage of messages with a scope basis, message count and preliminary exposure from the current sources.
Expand All @@ -20,6 +20,8 @@ The current slice is a responsive React/Vite local-first review workspace, a det
- Review reports and approved change requests can be exported as Markdown files without sending project data to an external service.
- Pilot safeguards include a 10 MB source limit, required scope/message source validation, source reclassification and removal, stable finding IDs, basic negation and sender-role handling, included/excluded scope matching, and multipart email/plain-text Slack fallbacks.
- The pilot visual system remains the original warm paper/ink/orange workspace design; the page header and demo banner have explicit layout styles so pilot-state additions do not fall back to browser defaults.
- An initial order email can be used as the scope source: order/request-style filenames are scope candidates, and manually reclassified prose emails produce scope items without requiring Markdown bullets.
- RTF order exports saved with an `.md` extension are normalized locally; cancellation/reply email filenames stay communication candidates even when their quoted history mentions order totals.

## Next technical slice

Expand Down
6 changes: 3 additions & 3 deletions docs/product-status.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,7 +30,7 @@ Finance and sales users are not expected to use GitHub, Node.js or npm. The host
## 2. How the current MVP works

```text
SOW / contract + client communication
SOW / contract / brief / initial order email + client communication
↓
Local file upload
↓
Expand All @@ -47,7 +47,7 @@ SOW / contract + client communication

The current workflow is:

1. Upload a scope document.
1. Upload a scope document or initial order email.
2. Upload a communication export.
3. Run the analysis.
4. Review potential scope-drift findings.
Expand All @@ -71,7 +71,7 @@ The current workflow is:

| Source | Current method | Status |
| --- | --- | --- |
| SOW / contract / brief | TXT, MD | Working |
| SOW / contract / brief / initial order email | TXT, MD (including common RTF-in-MD order exports), EML | Working |
| Gmail and other email services | EML or text export | Working through export |
| Slack | JSON or text export | Working through export |
| Telegram | JSON export | Working through export |
Expand Down
126 changes: 119 additions & 7 deletions src/analysis.ts
Original file line number Diff line number Diff line change
Expand Up @@ -147,12 +147,12 @@ export function validateSources(sources: SourceDocument[]): SourceValidation {
const messageSources = sources.filter((source) => source.kind === 'messages')
const unknownSources = sources.filter((source) => source.kind === 'unknown')

if (!scopeSources.length) errors.push('Add at least one scope document (SOW, contract or brief).')
if (!scopeSources.length) errors.push('Add at least one scope source (SOW, contract, brief or initial order email). If the agreed scope is in an email, classify that source as Scope document in the source list.')
if (!messageSources.length) errors.push('Add at least one communication export from Slack, email or WhatsApp.')
if (unknownSources.length) warnings.push(`Unclassified sources will not be analysed: ${unknownSources.map((source) => source.name).join(', ')}.`)

if (scopeSources.length && !scopeSources.some((source) => extractScopeItems(source.id, source.content).length)) {
errors.push('The scope document has no readable bullet points. Use headings with bullet or numbered deliverables.')
errors.push('The scope source has no readable deliverables. Use bullet points or a clear initial order email with the requested work.')
}
if (messageSources.length && !messageSources.some((source) => extractMessages(source).length)) {
errors.push('The communication export has no readable messages. Check the export format or reclassify the source.')
Expand Down Expand Up @@ -184,11 +184,19 @@ export function analyzeSources(sources: SourceDocument[]): AnalysisResult {
}

function inferKind(name: string, content: string): SourceKind {
const lowerName = name.toLowerCase()
const scopeNameSignal = /(^|[-_.])(sow|scope|contract|agreement|brief|proposal)([-_.]|$)/.test(lowerName)
const scopeContentSignal = /(^|\n)\s*#{1,6}\s*(included|excluded|scope|deliverables|assumptions|requirements)/im.test(content)
const lowerName = name.toLowerCase().replace(/\s+/g, '-')
const normalizedContent = normalizeDocumentText(content)
const replyOrCancellationNameSignal = /(^|[-_.])(re|reply|fwd|forward|thread|cancel|cancellation)([-_.]|$)/.test(lowerName)
const scopeNameSignal = !replyOrCancellationNameSignal
&& /(^|[-_.])(sow|scope|contract|agreement|brief|proposal|order|request|intake|kickoff|requirements?)([-_.]|$)/.test(lowerName)
const scopeContentSignal = !replyOrCancellationNameSignal && (
/(^|\n)\s*#{1,6}\s*(included|excluded|scope|deliverables|assumptions|requirements)/im.test(normalizedContent)
|| /\b(new order|order number|shipping cost|total cost|deliverables?)\b/i.test(normalizedContent)
)
const messageNameSignal = /slack|email|message|messenger|telegram|whatsapp|facebook|meta|thread|chat|linear|jira/.test(lowerName)
const messageContentSignal = looksLikeMessageExport(content) || /(^|\n)\s*(from|subject|date):/im.test(content)
const messageContentSignal = looksLikeMessageExport(content)
|| /(^|\n)\s*(from|subject|date):/im.test(content)
|| /^\s*(?:\[)?\d{4}[-/.]\d{1,2}[-/.]\d{1,2}[^|]*\|\s*[^|]+\|\s*.+$/m.test(content)

if (scopeNameSignal || scopeContentSignal) return 'scope'
if (messageNameSignal || messageContentSignal || detectPilotChannel(name, content)) return 'messages'
Expand Down Expand Up @@ -217,10 +225,12 @@ function getFormat(name: string): SourceFormat {
}

function extractScopeItems(sourceId: string, content: string): ScopeItem[] {
const normalizedContent = normalizeDocumentText(content)
const richText = isRichTextDocument(content)
let currentSection = 'Scope'
const occurrences = new Map<string, number>()

return content.split(/\r?\n/).flatMap((line) => {
const bulletItems = normalizedContent.split(/\r?\n/).flatMap((line) => {
const trimmed = line.trim()
if (!trimmed) return []
const heading = trimmed.match(/^#{1,6}\s*(.+)$/)
Expand All @@ -241,6 +251,108 @@ function extractScopeItems(sourceId: string, content: string): ScopeItem[] {
excluded: /excluded|no |not included|out of scope/i.test(currentSection) || /^(no |not included|excluded)/i.test(text),
}]
})

if (bulletItems.length) return bulletItems

const proseLines = richText
? normalizedContent.split(/\r?\n/)
: normalizedContent.replace(/\r\n/g, '\n').split(/\n\s*\n/).flatMap((block) => block.split(/(?<=[.!?])\s+(?=[A-Z0-9])/))

return proseLines
.map((line) => line.replace(/\s+/g, ' ').trim())
.filter((line) => line.length >= (richText ? 5 : 16))
.filter((line) => !/^(title|quantity|weight|cost)$/i.test(line))
.filter((line) => !/^(from|to|cc|bcc|subject|date|reply-to|mime-version|content-type|content-transfer-encoding|message-id):/i.test(line))
.filter((line) => !/^[-=]{3,}$/.test(line))
.filter((line) => !/^>/.test(line))
.filter((line) => !/^(hi|hello|dear|thanks|thank you|best|regards)[,!]?$/i.test(line))
.filter((line) => !/@/.test(line))
.map((text, index) => {
const key = normalizeForMatching(text)
const occurrence = occurrences.get(key) ?? 0
occurrences.set(key, occurrence + 1)
return {
id: `${sourceId}-scope-${stableHash(`${key}|${occurrence}|prose`)}`,
section: 'Initial order',
text,
excluded: /excluded|no |not included|out of scope/i.test(text),
}
})
}

function isRichTextDocument(content: string): boolean {
return /^\s*\{\\rtf/i.test(content)
}

function normalizeDocumentText(content: string): string {
if (!isRichTextDocument(content)) return content

const destinationWords = new Set([
'fonttbl', 'colortbl', 'stylesheet', 'info', 'generator', 'pict', 'object', 'filetbl',
'header', 'footer', 'listtable', 'listoverridetable', 'themedata', 'xmlnstbl', 'datastore',
])
const output: string[] = []
const stack: boolean[] = []
let skipGroup = false
let unicodeFallback = 1
let skipFallback = 0

for (let index = 0; index < content.length; index += 1) {
const character = content[index]
if (character === '{') {
stack.push(skipGroup)
continue
}
if (character === '}') {
skipGroup = stack.pop() ?? false
continue
}
if (character !== '\\') {
if (skipFallback > 0) {
skipFallback -= 1
} else if (!skipGroup) {
output.push(character)
}
continue
}

const next = content[index + 1]
if (next === '\\' || next === '{' || next === '}' || next === '~' || next === '-' || next === '_') {
if (!skipGroup && skipFallback === 0) output.push(next === '~' ? ' ' : next)
index += 1
continue
}
if (next === "'") {
const hex = content.slice(index + 2, index + 4)
if (/^[0-9a-f]{2}$/i.test(hex)) {
if (!skipGroup && skipFallback === 0) output.push(String.fromCharCode(Number.parseInt(hex, 16)))
index += 3
continue
}
}

const control = content.slice(index + 1).match(/^([a-z]+)(-?\d+)? ?/i)
if (!control) {
if (next === '*') skipGroup = true
index += 1
continue
}

const word = control[1].toLowerCase()
const parameter = control[2] ? Number(control[2]) : undefined
index += control[0].length
if (destinationWords.has(word)) skipGroup = true
if (word === 'uc' && parameter !== undefined) unicodeFallback = Math.max(0, parameter)
if (word === 'u' && parameter !== undefined) {
const codePoint = parameter < 0 ? parameter + 65536 : parameter
if (!skipGroup) output.push(String.fromCharCode(codePoint))
skipFallback = unicodeFallback
}
if (!skipGroup && ['par', 'line', 'cell', 'row'].includes(word)) output.push('\n')
if (!skipGroup && word === 'tab') output.push('\t')
}

return output.join('').replace(/\u00a0/g, ' ').replace(/[ \t]+\n/g, '\n').replace(/\n{3,}/g, '\n\n').trim()
}

function extractMessages(source: SourceDocument): MessageRecord[] {
Expand Down
1 change: 1 addition & 0 deletions src/main.tsx
Original file line number Diff line number Diff line change
Expand Up @@ -102,6 +102,7 @@ function App() {

const reclassifySource = (id: string, kind: SourceDocument['kind']) => {
setSources((current) => current.map((source) => source.id === id ? { ...source, kind } : source))
setSourceError(null)
setAnalysisReady(false)
setReviewingId(null)
}
Expand Down
59 changes: 58 additions & 1 deletion tests/connectors.test.ts
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
import { describe, expect, it } from 'vitest'
import { analyzeSources, demoSources, validateSources } from '../src/analysis'
import { analyzeSources, demoSources, parseSourceFile, validateSources } from '../src/analysis'
import { detectPilotChannel, parsePilotExport } from '../src/connectors'

describe('pilot channel adapters', () => {
Expand Down Expand Up @@ -140,4 +140,61 @@ describe('scope analysis', () => {

expect(validation.errors).toContain('Add at least one communication export from Slack, email or WhatsApp.')
})

it('accepts a prose initial order email as the scope source', () => {
const sources = [
{
id: 'initial-order',
name: 'initial-order.eml',
kind: 'scope' as const,
format: 'eml' as const,
content: [
'From: client@example.com',
'Subject: Website order',
'',
'We need a public marketing site with a responsive mobile layout.',
].join('\n'),
},
{
...demoSources[1],
content: JSON.stringify([{ user: 'Client', text: 'Can we add a partner dashboard?' }]),
},
]

expect(validateSources(sources).errors).toEqual([])
expect(analyzeSources(sources).scopeItemsCount).toBe(1)
})

it('classifies an initial order filename as a scope candidate', async () => {
const source = await parseSourceFile(new File([
'From: client@example.com\nSubject: Website order\n\nWe need a public marketing site.',
], 'initial-order.eml'))

expect(source.kind).toBe('scope')
})

it('classifies a structured conversation text export as messages', async () => {
const source = await parseSourceFile(new File([
'2026-08-03 10:00 | Client | Can we add a dashboard?',
], 'conversation.txt'))

expect(source.kind).toBe('messages')
})

it('classifies and reads an RTF order export as a scope source', async () => {
const source = await parseSourceFile(new File([
'{\\rtf1\\ansi\\uc0 This is a new order\\par Product: Trail rack\\par Quantity: 1\\par Total cost: 100 CAD}',
], '2607-22.md'))

expect(source.kind).toBe('scope')
expect(analyzeSources([source, demoSources[1]]).scopeItemsCount).toBeGreaterThan(0)
})

it('keeps a cancellation reply email in communications', async () => {
const source = await parseSourceFile(new File([
'From: seller@example.com\nSubject: RE: order cancellation\nDate: Sun, 02 Aug 2026\n\nThe order will not proceed. The quoted history mentions a shipping cost and total cost.',
], 'RE_order_cancellation.eml'))

expect(source.kind).toBe('messages')
})
})
Loading