forked from lentrekin1/devmind
-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy path.cursorrules
More file actions
365 lines (285 loc) · 16.3 KB
/
Copy path.cursorrules
File metadata and controls
365 lines (285 loc) · 16.3 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
# Best Practices
typescript:
- use TypeScript for all code
- prefer interfaces over types
- avoid enums, use maps
nextjs:
- rely on Next.js App Router for state management
- prioritize Web Vitals (LCP, CLS, FID)
react:
- use functional components
- favor server components for better performance
- use 'use client' only for essential Web API access
tailwind:
- implement responsive design
- use mobile-first approach
- combine Tailwind utility classes with component-specific styles
- components should fill the screen horizontally unless otherwise specified
security:
- use Clerk for authentication handling
- implement server-side security best practices with SQLAlchemy
performance:
- minimize 'useEffect' and 'useState' in React components
- optimize images with WebP format and lazy loading
database:
- use Prisma to connect from Next.js to Postgres
- use SQLAlchemy for backend database interactions
- ensure secure database connections
build:
- use pnpm for package management
- deploy with Vercel for serverless hosting
architecture:
- website code should not access database directly
- website code should use a nextjs server action (in /actions) that then calls a database function (in /lib/db) to get data
- server actions should be called xxxAction and the file name should be xxx.ts
- server actions should generally use the withAuth method in /lib/auth
- server actions never access the database directly, only through database methods in lib/db!!
- avoid unnecessary wrappers
- always try to follow existing patterns in the codebase. Look for such patterns before writing code.
project description:
- the project is a next.js website where users upload documents (grouped by client), make a questionnaire, and then use AI to answer those questions for each client, finding relevant evidence in the documents
- always use best practices
# Instructions
You are a multi-agent system coordinator, playing two roles in this environment: Planner and Executor. You will decide the next steps based on the current state of `Multi-Agent Scratchpad` section in the `.cursorrules` file. Your goal is to complete the user's (or business's) final requirements. The specific instructions are as follows:
## Role Descriptions
1. Planner
* Responsibilities: Perform high-level analysis, break down tasks, define success criteria, evaluate current progress. When doing planning, always use high-intelligence models (OpenAI o1 via `tools/plan_exec_llm.py`). Don't rely on your own capabilities to do the planning.
* Actions: Invoke the Planner by calling `venv/bin/python tools/plan_exec_llm.py --prompt {any prompt}`. You can also include content from a specific file in the analysis by using the `--file` option: `venv/bin/python tools/plan_exec_llm.py --prompt {any prompt} --file {path/to/file}`. It will print out a plan on how to revise the `.cursorrules` file. You then need to actually do the changes to the file. And then reread the file to see what's the next step.
2) Executor
* Responsibilities: Execute specific tasks instructed by the Planner, such as writing code, running tests, handling implementation details, etc.. The key is you need to report progress or raise questions to the Planner at the right time, e.g. after completion some milestone or after you've hit a blocker.
* Actions: When you complete a subtask or need assistance/more information, also make incremental writes or modifications to the `Multi-Agent Scratchpad` section in the `.cursorrules` file; update the "Current Status / Progress Tracking" and "Executor's Feedback or Assistance Requests" sections. And then change to the Planner role.
## Document Conventions
* The `Multi-Agent Scratchpad` section in the `.cursorrules` file is divided into several sections as per the above structure. Please do not arbitrarily change the titles to avoid affecting subsequent reading.
* Sections like "Background and Motivation" and "Key Challenges and Analysis" are generally established by the Planner initially and gradually appended during task progress.
* "Current Status / Progress Tracking" and "Executor's Feedback or Assistance Requests" are mainly filled by the Executor, with the Planner reviewing and supplementing as needed.
* "Next Steps and Action Items" mainly contains specific execution steps written by the Planner for the Executor.
## Workflow Guidelines
* After you receive an initial prompt for a new task, update the "Background and Motivation" section, and then invoke the Planner to do the planning.
* When thinking as a Planner, always use the local command line `python tools/plan_exec_llm.py --prompt {any prompt}` to call the o1 model for deep analysis, recording results in sections like "Key Challenges and Analysis" or "High-level Task Breakdown". Also update the "Background and Motivation" section.
* When you as an Executor receive new instructions, use the existing cursor tools and workflow to execute those tasks. After completion, write back to the "Current Status / Progress Tracking" and "Executor's Feedback or Assistance Requests" sections in the `Multi-Agent Scratchpad`.
* If unclear whether Planner or Executor is speaking, declare your current role in the output prompt.
* Continue the cycle unless the Planner explicitly indicates the entire project is complete or stopped. Communication between Planner and Executor is conducted through writing to or modifying the `Multi-Agent Scratchpad` section.
Please note:
* Note the task completion should only be announced by the Planner, not the Executor. If the Executor thinks the task is done, it should ask the Planner for confirmation. Then the Planner needs to do some cross-checking.
* Avoid rewriting the entire document unless necessary;
* Avoid deleting records left by other roles; you can append new paragraphs or mark old paragraphs as outdated;
* When new external information is needed, you can use command line tools (like search_engine.py, llm_api.py), but document the purpose and results of such requests;
* Before executing any large-scale changes or critical functionality, the Executor should first notify the Planner in "Executor's Feedback or Assistance Requests" to ensure everyone understands the consequences.
* During you interaction with the user, if you find anything reusable in this project (e.g. version of a library, model name), especially about a fix to a mistake you made or a correction you received, you should take note in the `Lessons` section in the `.cursorrules` file so you will not make the same mistake again.
# Tools
Note all the tools are in python. So in the case you need to do batch processing, you can always consult the python files and write your own script.
## Screenshot Verification
The screenshot verification workflow allows you to capture screenshots of web pages and verify their appearance using LLMs. The following tools are available:
1. Screenshot Capture:
```bash
venv/bin/python tools/screenshot_utils.py URL [--output OUTPUT] [--width WIDTH] [--height HEIGHT]
```
2. LLM Verification with Images:
```bash
venv/bin/python tools/llm_api.py --prompt "Your verification question" --provider {openai|anthropic} --image path/to/screenshot.png
```
Example workflow:
```python
from screenshot_utils import take_screenshot_sync
from llm_api import query_llm
# Take a screenshot
screenshot_path = take_screenshot_sync('https://example.com', 'screenshot.png')
# Verify with LLM
response = query_llm(
"What is the background color and title of this webpage?",
provider="openai", # or "anthropic"
image_path=screenshot_path
)
print(response)
```
## LLM
You always have an LLM at your side to help you with the task. For simple tasks, you could invoke the LLM by running the following command:
```
venv/bin/python ./tools/llm_api.py --prompt "What is the capital of France?" --provider "anthropic"
```
The LLM API supports multiple providers:
- OpenAI (default, model: gpt-4o)
- Azure OpenAI (model: configured via AZURE_OPENAI_MODEL_DEPLOYMENT in .env file, defaults to gpt-4o-ms)
- DeepSeek (model: deepseek-chat)
- Anthropic (model: claude-3-sonnet-20240229)
- Gemini (model: gemini-pro)
- Local LLM (model: Qwen/Qwen2.5-32B-Instruct-AWQ)
But usually it's a better idea to check the content of the file and use the APIs in the `tools/llm_api.py` file to invoke the LLM if needed.
## Web browser
You could use the `tools/web_scraper.py` file to scrape the web.
```
venv/bin/python ./tools/web_scraper.py --max-concurrent 3 URL1 URL2 URL3
```
This will output the content of the web pages.
## Search engine
You could use the `tools/search_engine.py` file to search the web.
```
venv/bin/python ./tools/search_engine.py "your search keywords"
```
This will output the search results in the following format:
```
URL: https://example.com
Title: This is the title of the search result
Snippet: This is a snippet of the search result
```
If needed, you can further use the `web_scraper.py` file to scrape the web page content.
# Lessons
## User Specified Lessons
- You have a python venv in ./venv. Use it.
- Include info useful for debugging in the program output.
- Read the file before you try to edit it.
- Due to Cursor's limit, when you use `git` and `gh` and need to submit a multiline commit message, first write the message in a file, and then use `git commit -F <filename>` or similar command to commit. And then remove the file. Include "[Cursor] " in the commit message and PR title.
## Cursor learned
- For search results, ensure proper handling of different character encodings (UTF-8) for international queries
- Add debug information to stderr while keeping the main output clean in stdout for better pipeline integration
- When using seaborn styles in matplotlib, use 'seaborn-v0_8' instead of 'seaborn' as the style name due to recent seaborn version changes
- Use `gpt-4o` as the model name for OpenAI. It is the latest GPT model and has vision capabilities as well. `o1` is the most advanced and expensive model from OpenAI. Use it when you need to do reasoning, planning, or get blocked.
- Use `claude-3-5-sonnet-20241022` as the model name for Claude. It is the latest Claude model and has vision capabilities as well.
# Multi-Agent Scratchpad
## Background and Motivation
The project needs a new timeline feature that allows users to visualize and interact with client records chronologically. The key focus is on using AI/LLMs to create highly accurate, detailed timelines from document content, leveraging the existing OCR and vector embeddings infrastructure. The timeline should be both viewable in the web interface and exportable to PDF, similar to the existing report functionality. This will help legal professionals understand the sequence of events, identify key moments, and discover relationships between different documents in their cases.
## Key Challenges and Analysis
1. AI/LLM Processing Challenges:
- Need to extract multiple types of temporal information from OCR text:
* Explicit dates (e.g., "January 1, 2024")
* Relative dates (e.g., "two weeks after the incident")
* Event sequences (e.g., "following the initial consultation")
* Date ranges (e.g., "throughout 2023")
- Must understand document context to identify relevant dates
- Need to handle conflicting dates within documents
- Must determine event importance/relevance
- Should leverage existing vector embeddings for semantic similarity
2. Data Quality Challenges:
- Need to handle OCR errors in date recognition
- Documents may contain ambiguous dates
- Need to validate extracted dates against case context
- Must handle different date formats and standards
3. Timeline Construction Challenges:
- Need to establish causal relationships between events
- Must identify and group related events
- Should detect and highlight key milestone events
- Need to handle overlapping or concurrent events
- Must generate both interactive web view and static PDF export
## AI Processing Strategy
1. Initial Processing:
- Leverage existing OCR text from database
- Use vector embeddings to find semantically related events
- Identify document type and context
- Extract metadata (e.g., document creation date)
2. Date Extraction (Using Claude 3.5 Sonnet):
- First pass: Extract all potential dates and temporal references from OCR text
- Second pass: Analyze context around each date
- Third pass: Validate dates against document metadata
- Store confidence scores for each extracted date
- Use vector similarity to find related events across documents
3. Event Analysis (Using GPT-4):
- Identify event descriptions associated with dates
- Classify event types and importance
- Extract key entities and relationships
- Determine event duration (point-in-time vs. period)
- Use embeddings to cluster similar events
4. Timeline Construction (Using Claude 3.5 Sonnet):
- Establish chronological sequence
- Identify causal relationships between events
- Group related events using semantic similarity
- Generate event summaries
- Assign confidence scores to timeline entries
- Create hierarchical event structure for different detail levels
5. Quality Assurance:
- Cross-validate dates between related documents
- Check for logical consistency in event sequence
- Identify and flag potential conflicts or anomalies
- Allow user verification and correction
- Validate timeline coherence before export
## Export Processing
1. PDF Generation:
- Create both detailed and summary timeline views
- Include confidence scores and source citations
- Generate event relationship diagrams
- Add table of contents and navigation
- Support different timeline layouts (chronological, grouped, milestone-focused)
2. Export Features:
- Customizable detail levels
- Include/exclude specific event types
- Highlight key events
- Source document references
- Interactive elements in PDF (e.g., clickable citations)
## Verifiable Success Criteria
1. Extraction Accuracy:
- >95% accuracy for explicit dates
- >90% accuracy for relative dates
- >85% accuracy for event relationships
- <1% false positive rate for date extraction
2. Timeline Quality:
- Correct chronological ordering
- Meaningful event grouping using semantic similarity
- Clear indication of confidence levels
- Proper handling of uncertain dates
- Accurate source citations
3. User Experience:
- Clear visualization of event relationships
- Easy verification of AI-extracted dates
- Simple correction mechanism for errors
- Useful event summaries and context
- Seamless export to PDF
## High-level Task Breakdown
Phase 1: AI Pipeline Development
1. Enhance existing Modal processing
- Add date extraction to processing pipeline
- Implement multi-pass date validation
- Create event analysis pipeline
- Build timeline construction logic
- Integrate with vector embeddings
2. Create AI prompts and templates
- Date extraction prompts
- Event analysis prompts
- Relationship detection prompts
- Timeline construction prompts
- PDF generation prompts
3. Implement validation and QA
- Cross-document validation
- Consistency checking
- Confidence scoring
- Error detection
- Export validation
Phase 2: Data Storage & API
1. Create data models
- Timeline events
- Event relationships
- Confidence scores
- User corrections
- Export preferences
2. Implement API endpoints
- Timeline generation
- Event management
- Validation endpoints
- Correction handling
- Export generation
Phase 3: Frontend & Export Development
1. Timeline visualization
- Interactive web view
- Relationship indicators
- Confidence level display
- Export options interface
- PDF preview
## Current Status / Progress Tracking
Status: Planning Phase
Next up: Begin Phase 1 - AI Pipeline Development
## Next Steps and Action Items
1. Design AI processing pipeline
- Create prompt templates
- Define validation rules
- Plan QA process
- Design export formats
2. Set up testing framework
- Use existing OCR'd documents
- Create test cases
- Define accuracy metrics
- Test export quality
3. Begin prompt engineering
- Date extraction prompts
- Event analysis prompts
- Timeline construction prompts
- PDF generation prompts
## Executor's Feedback or Assistance Requests
Awaiting approval of AI-focused approach and prompt engineering strategy before proceeding with implementation. Need to confirm if any specific PDF export requirements or limitations need to be considered.