Skip to content

Latest commit

 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

The Motivation / Use-case:

Problem Statement: Every month, my daughter's shool Parent-Teacher Council (PTC), just like every other school in USA, meets to discuss important issues related to kids, their leraning, teachers, the school environment, budgetting, etc. As a secretary to the board, my job is to keep the minutes of the meeting, and get it approved in the next meeting. You would think keeping meeting minutes is such a trivial job now a days with Zoom, Google-Meet, etc, supporting AI powered meeting minutes generation, but the School & the School Districts are not so fast in embracing these newer technologies due to governance, legacy subscriptions, and budgetting constraints. In our case, the School District has a legacy Zoom subscription and a policy that does not allow automatic meeting minutes.

Solution: Build an Agentic AI sytem that, given a meeting recording, generates meeting minutes, creates a document of preferred type, and stores or uploads it at a preferred location.

Why Agentic AI? While at first glance this might seem like a glorified automation, and in its current form that's somewhat true, the potential to make this more flexible by leverging the skills of the LLMs exists: (i) the encoding format of the input audio file need not be fixed, let the LLM figure out how to deal with any format; (ii) the file format for the MoMs need not be fixed to google doc, let the LLM figure out how to generate if other formats are asked for (potentially LLM generating the code on the fly with appropriate prompt changes TBD); (iii) the upload destination can be anything (School PTC Website, Box, AWS S3, ...); let LLM figure out what tools to use, and potentiall come up with the required code as required (TBD). Hence an Agentic AI system with differnt agents for each task is deliberately chosen for these future enhancements.

Why Input is a JSON File: The longer term goal is to expose an API that that takes the required arguments and does the job, so this tool can be easily integrated with the school's PTC Website, but this is a future enhancement.

PTC Meeting Minutes — CrewAI Pipeline

An agentic AI pipeline that takes a Zoom recording (.mp4 / .m4a), transcribes it, generates structured meeting minutes, and publishes them to Google Drive — fully automated, end-to-end.

Built with Python · CrewAI · OpenAI Whisper · Google Docs & Drive APIs.
Runs in Jupyter Notebook or Google Colab.


Pipeline Architecture

  ┌─────────────────────────────────────────────────────────────────────────────┐
  │                        PTC MoMs  —  Agent Pipeline                          │
  └─────────────────────────────────────────────────────────────────────────────┘

        input.json
   ┌──────────────────────┐
   │  media_file_path     │   .mp4 or .m4a Zoom recording
   │  meeting_date        │   YYYY-MM-DD string
   │  drive_folder_id     │   Google Drive destination folder ID
   │  credentials_path    │   OAuth2 credentials JSON path
   └──────────┬───────────┘
              │
              ▼
   ╔══════════════════════╗
   ║  1. Orchestrator     ║  • Loads and parses input.json
   ║     Agent            ║  • Validates every field (file existence,
   ╚══════════╦═══════════╝    extension, date format, credentials)
              ║ validated config
              ▼
   ╔══════════════════════╗
   ║  2. Transcript       ║  • Extracts audio from .mp4 if needed
   ║     Agent            ║  • Runs OpenAI Whisper locally (CPU/GPU)
   ╚══════════╦═══════════╝  • Returns full verbatim transcript
              ║ verbatim transcript text
              ▼
   ╔══════════════════════╗
   ║  3. Secretary        ║  • Reads transcript with an LLM
   ║     Agent            ║  • Produces six structured sections:
   ╚══════════╦═══════════╝    Executive Summary (≤50 words)
              ║                Motions Passed
              ║                Action Items  (person → task)
              ║                Decisions Taken
              ║                Open Issues
              ║                Key Messages
              ║ markdown meeting minutes
              ▼
   ╔══════════════════════╗
   ║  4. Google Docs      ║  • Creates a new Google Document
   ║     Creator Agent    ║  • Applies HEADING_1 / HEADING_2 styles
   ╚══════════╦═══════════╝  • Returns doc ID and shareable URL
              ║ doc_id + doc_url
              ▼
   ╔══════════════════════╗
   ║  5. Drive Uploader   ║  • Moves the doc into the specified
   ║     Agent            ║    Google Drive folder
   ╚══════════╦═══════════╝  • Returns confirmation + shareable link
              ║
              ▼
   ✅  Meeting Minutes live in Google Drive

Each stage is also a standalone notebook (notebooks/0N_*.ipynb) that can be run and debugged independently.


Project Structure

PTC-MoMs-CrewAI/
│
├── main_pipeline.ipynb              ← ▶  Run this for the full end-to-end pipeline
│
├── notebooks/
│   ├── 01_orchestrator.ipynb        ← Stage 1 – input validation
│   ├── 02_transcript_agent.ipynb    ← Stage 2 – Whisper transcription
│   ├── 03_secretary_agent.ipynb     ← Stage 3 – meeting minutes (LLM)
│   ├── 04_docs_creator_agent.ipynb  ← Stage 4 – Google Docs creation
│   └── 05_drive_uploader_agent.ipynb← Stage 5 – Google Drive upload
│
├── src/
│   ├── config.py                    ← Pydantic input validation model
│   ├── llm_config.py                ← LLM factory (swap providers in .env)
│   ├── state.py                     ← JSON-backed state shared across notebooks
│   └── tools/
│       ├── transcription_tool.py    ← CrewAI tool wrapping Whisper
│       ├── gdocs_tool.py            ← CrewAI tool wrapping Google Docs API
│       └── gdrive_tool.py           ← CrewAI tool wrapping Google Drive API
│
├── requirements.txt
├── .env.example                     ← Copy → .env and fill in your keys
├── input_template.json              ← Copy → input.json and fill in your paths
└── pipeline_state.json              ← Auto-created at runtime (inter-notebook state)

Prerequisites

Requirement Notes
Python 3.10+ 3.11 recommended
ffmpeg Required for .mp4 audio extraction. Install via your OS package manager.
OpenAI API key Or any supported LLM provider (see Switching LLM)
Google Cloud project With Google Docs API and Google Drive API enabled
OAuth2 credentials JSON Downloaded from Google Cloud Console

Install ffmpeg

# macOS
brew install ffmpeg

# Ubuntu / Debian
sudo apt install ffmpeg

# Windows — download from https://ffmpeg.org/download.html

Enable Google APIs

  1. Go to Google Cloud Console → APIs & Services → Library
  2. Enable Google Docs API
  3. Enable Google Drive API
  4. Go to APIs & Services → Credentials → Create Credentials → OAuth client ID
  5. Choose Desktop app, download the JSON → this is your credentials.json

Installation

Why a dedicated venv?
crewai supports Python 3.10–3.13 only. If JupyterLab is installed via Homebrew
(or any other method that bundles its own Python), its kernel may be Python 3.14+
and will fail to install crewai. The steps below create a Python 3.11 venv and
register it as a named Jupyter kernel so JupyterLab uses the right interpreter.

# 1. Clone the repo (or copy the folder)
git clone <repo-url>
cd PTC-MoMs-CrewAI

# 2. Create a venv with Python 3.11 (binary name on macOS Homebrew is python3.11)
python3.11 -m venv venv
source venv/bin/activate          # Windows: venv\Scripts\activate

# 3. Install all dependencies (including ipykernel)
pip install -r requirements.txt

# 4. Register this venv as a Jupyter kernel
python -m ipykernel install --user \
    --name ptc-moms-py311 \
    --display-name "PTC MoMs (Python 3.11)"

# 5. Launch JupyterLab (can be the system/Homebrew one — kernel is separate)
jupyter lab

Once JupyterLab opens:

  1. Open main_pipeline.ipynb
  2. In the top-right corner click the kernel name → Change Kernel
  3. Select "PTC MoMs (Python 3.11)"
  4. Do the same for any individual notebook you open

The first code cell in every notebook verifies the Python version at runtime and
raises a clear error if the wrong kernel is selected.

Google Colab: Skip steps 1–5. The first cell in every notebook runs
pip install -r requirements.txt automatically, and Colab's Python (3.10) is compatible.


Configuration

Step 1 — Create .env

cp .env.example .env

Open .env and set your LLM API key:

# .env
LLM_PROVIDER=openai
LLM_MODEL=gpt-4o
LLM_TEMPERATURE=0.1

OPENAI_API_KEY=sk-...          # ← required for default OpenAI setup

# Whisper settings (see Transcription section below)
WHISPER_MODEL_SIZE=base
WHISPER_DEVICE=cpu

Step 2 — Create input.json

cp input_template.json input.json

Edit input.json with your meeting details:

{
  "media_file_path":        "/absolute/path/to/zoom_recording.mp4",
  "meeting_date":           "2024-04-15",
  "google_drive_folder_id": "1BxiMVs0XRA5nFMdKvBdBZjgmUUqptlbs",
  "oauth2_credentials_path":"/absolute/path/to/credentials.json"
}

Input parameter reference

Parameter Type Description
media_file_path string Absolute path to the Zoom recording. Supported: .mp4, .m4a
meeting_date string Date the meeting was held. Format: YYYY-MM-DD
google_drive_folder_id string The folder ID from the Google Drive URL: drive.google.com/drive/folders/<ID>
oauth2_credentials_path string Absolute path to the OAuth2 credentials.json downloaded from Google Cloud

Finding your Drive folder ID:
Open the folder in Google Drive. The URL looks like
https://drive.google.com/drive/folders/1BxiMVs0XRA5nFMdKvBdBZjgmUUqptlbs
The long alphanumeric string at the end is the folder ID.


Running the Pipeline

Option A — Full pipeline (recommended)

Open main_pipeline.ipynb in Jupyter Lab / Jupyter Notebook / Google Colab and run all cells top-to-bottom.

# Local Jupyter
jupyter lab main_pipeline.ipynb

The notebook will:

  1. Install dependencies
  2. Load and validate input.json
  3. Transcribe the recording (may take several minutes for long files)
  4. Generate meeting minutes
  5. Create and upload the Google Doc
  6. Print the final shareable URL

First run — Google auth: A browser window will open asking you to grant
access to Google Docs and Drive. After consent, a token.json is saved next
to your credentials.json so subsequent runs skip this step.
In Colab, authentication is handled automatically via google.colab.auth.

Option B — Run individual notebooks

Each notebooks/0N_*.ipynb notebook can run on its own. Notebooks exchange data via pipeline_state.json (created automatically).

Run in order:                        Or jump to any stage by filling in the
                                     OVERRIDE_* variables at the top of the notebook.
01_orchestrator.ipynb   ──┐
02_transcript_agent.ipynb ├─► pipeline_state.json ──► each notebook reads/writes here
03_secretary_agent.ipynb  │
04_docs_creator_agent.ipynb
05_drive_uploader_agent.ipynb

Standalone example — run only the Secretary agent on an existing transcript:

Open 03_secretary_agent.ipynb and set:

# Near the top of the notebook
OVERRIDE_TRANSCRIPT = """
Good morning everyone. Let us call the meeting to order...
[paste your transcript here]
"""

Then run all cells.


Switching LLM Provider

Change two lines in .env — no code changes required.

# Anthropic Claude
LLM_PROVIDER=anthropic
LLM_MODEL=claude-3-5-sonnet-20241022
ANTHROPIC_API_KEY=sk-ant-...

# Groq (fast inference)
LLM_PROVIDER=groq
LLM_MODEL=llama3-70b-8192
GROQ_API_KEY=gsk_...

# Ollama (local, no API key needed)
LLM_PROVIDER=ollama
LLM_MODEL=llama3

Supported providers are defined in src/llm_config.py. Adding a new one is a two-line change in that file.


Transcription Settings (faster-whisper)

The transcription agent uses faster-whisper running locally — no audio leaves your machine.
It uses CTranslate2 (not PyTorch/numba), so there is no LLVM dependency and it installs cleanly on all platforms.

WHISPER_MODEL_SIZE Speed Accuracy RAM (CPU)
tiny fastest lowest ~400 MB
base fast good ~500 MB
small medium better ~1 GB
medium slow great ~2.5 GB
large-v2 slowest best ~5 GB
large-v3 slowest best+ ~5 GB
# .env — CPU (default, no GPU required)
WHISPER_MODEL_SIZE=base
WHISPER_DEVICE=cpu
# WHISPER_COMPUTE_TYPE=int8     ← auto-set for CPU; override if needed

# .env — GPU (NVIDIA CUDA)
WHISPER_MODEL_SIZE=medium
WHISPER_DEVICE=cuda
# WHISPER_COMPUTE_TYPE=float16  ← auto-set for CUDA; override if needed

For .mp4 files, audio is extracted automatically to a temporary .wav file using moviepy + ffmpeg, then deleted after transcription.


Meeting Minutes Format

The Secretary agent always produces six sections in this order:

Section Description
Executive Summary ≤ 50 words — high-level overview of the meeting
Motions Passed Formally moved and seconded resolutions
Action Items [Person / Team]: [what they must do]
Decisions Taken Agreements reached during the meeting
Open Issues Topics raised but left unresolved or deferred
Key Messages Important statements and announcements from speakers

Troubleshooting

Symptom Fix
crewai>=0.80.0 — No matching distribution found Wrong kernel (Python 3.14+). Switch to "PTC MoMs (Python 3.11)" — see Installation above
Python X.Y is not supported (cell 1 error) Same — select the correct kernel in JupyterLab before running
Kernel "PTC MoMs (Python 3.11)" not listed Run python -m ipykernel install --user --name ptc-moms-py311 --display-name "PTC MoMs (Python 3.11)" inside the activated venv
FileNotFoundError: credentials.json Check the oauth2_credentials_path in input.json
Google Docs API has not been used… Enable the API in Google Cloud Console
Token has expired Delete token.json (next to credentials.json) and re-run
Whisper runs very slowly Use a smaller model (tiny or base) or set WHISPER_DEVICE=cuda
moviepy import error Run pip install moviepy or ensure ffmpeg is on your PATH
Empty transcript File may be silent or corrupt — verify playback before running

License

MIT

About

Agentic Code for PTC Meeting Minutes

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages