Institute for Social Research · University of Michigan

Applied AI for social research.

We put AI to work for social science, from statistical modeling to language-model agents and the automations in between. Bring us a problem; leave with a plan.

About us

A small team with a wide toolkit.

AI@ISR is the Institute for Social Research's applied AI group, continuing the work of ISR Data Science. We take on scoped engagements for research teams, sized to a single meeting or a full build. Four things we do:

SVC_101

Applied AI consulting

Where AI fits your study, and where it does not. We scope the question, pick the method, check the data-sensitivity and IRB fit, and leave you with a plan that will survive peer review, whether it takes one meeting or a full build.

SVC_102

AI agents and automation

Workflows that run themselves, from scripted pipelines to language-model agents that plan a multi-step task, call your tools, and stop for a human before anything final. Data intake, cleaning, coding, and reporting, with an audit trail of every step.

SVC_103

Modeling and machine learning

Statistical and machine learning models built for inference, not just prediction scores: regression to ensembles, causal and predictive designs, interpretable by default, and validated honestly before anyone cites them.

SVC_104

Language models on research data

Large language models applied to text, audio, and documents: coding open-ended responses, retrieval over document collections, transcription and redaction, and custom assistants. Open-weight models run on U-M infrastructure so sensitive data stays inside the university, and every output ships with an evaluation you can defend.

The team · Book a consult

Meet the team, pick a time.

We work as a team: book either of us and you get both sets of eyes on your problem. Choose based on who you've talked to or who fits the topic; we'll sort the rest. Consults are live over Google Meet or in person at ISR. Booking confirms instantly on both calendars, and you can reschedule from the invite.

Alexis Castellanos

ISR AI Program Manager

Alexis is an AI Program Manager with a background in Computer Science, specializing in generative AI applications, traditional machine learning, and intelligent workflow automation. They build technical solutions that bridge the gap between complex AI capabilities and practical, real-world utility. Working with research units across the University of Michigan, including CSCAR, Advanced Research Computing (ARC), and the Institute for Research on Innovation & Science (IRIS), Alexis helps scholars integrate modern data tools, streamline technical workflows, and accelerate academic discovery.

Book Alexis for new projects, scoping, and partnerships. Best first stop if we haven't worked together yet.

APPOINTMENT SCHEDULE EMBED · ALEXIS

Prefer a full page? Open Alexis's booking page in a new tab [link pending]

Loren Fang

AI Applications Developer

Loren leads technical execution and software development for ISR's AI initiatives, taking projects from concept and prototype to final delivery. Working directly with campus researchers, she handles hands-on implementation across the stack, from deploying open-source LLMs on HPC clusters to building custom automation pipelines. She focuses on transforming complex research needs into reliable, production-ready software tools.

Book Loren for working sessions on active projects and hands-on build help.

APPOINTMENT SCHEDULE EMBED · LOREN

Prefer a full page? Open Loren's booking page in a new tab [link pending]

// We keep a record of bookings (date, consultant, and your email) to manage appointments and improve the service.

Software tools

Research software, run on U-M infrastructure.

The ISR Software Toolkit is a set of services we built for qualitative research data. Jobs run on GPU nodes of the Great Lakes cluster, so recordings and transcripts stay inside the university: no third-party API, no vendor account. Two services are in pilot with research teams now. Read what each one does, then reach out and we will walk you through using it on your study.

TOOL_01 Pilot

Transcription

Audio or video in, a speaker-separated transcript out. Every turn carries a timestamp and a speaker label, so you know who said what, which a Zoom transcript does not tell you.

What it does

  • Separates and labels each speaker throughout the recording (Whisper large-v3 with PyAnnote diarization).
  • Detects the language automatically, or uses the one you set. English and Spanish interviews have been run for outside studies.
  • Puts a name on a known voice: enrol an interviewer from a short voice print and their turns carry the ID you choose. Nothing is stored between jobs.
  • Handles one speaker, two, or an auto-detected count. Turns it cannot attribute are marked, never merged into a neighbour.
  • Reads from your Mac, a mounted ISR drive, a synced Dropbox folder, or cluster storage in place.
  • Never overwrites: a repeat run is versioned beside the first, so a transcript under review is never replaced.

What you get

Formats
Speaker turns as .txt, plus .vtt, .srt, and .json on request, delivered to a dated folder beside your audio.
Summary
An optional file with speaker count, turns, duration, and speaking time per speaker, and no transcript text, for checking an interview went as reported.
Options
Timestamps as hh:mm:ss or seconds. Subtitle cues per sentence or per speaker turn.
Speed
About two minutes of GPU time for a 25-minute interview, plus cluster queue time.

// Transcripts are a first pass for human correction, not a finished record.

TOOL_02 Pilot

PII Redaction

Replaces identifying details in transcripts and open-ended survey responses with category tags, so the text stays readable and codable while the person in it is not identifiable.

What it does

  • Tags 11 categories: names, research team members, relationships, ages, locations, organizations, personal characteristics, health conditions, contact details, ID numbers, and financial details. Dates on request, per HIPAA Safe Harbor.
  • Works in three layers: pattern matching for structured identifiers (emails, phones, SSNs, card numbers), named-entity recognition for places and dates, and a language model on the cluster for anything that needs context.
  • The pattern layer is a floor. If the model fails, structured identifiers are still removed.
  • Precision first: it matches only what it is sure of, so your research data is not eaten. A short study description tells the model what counts as data in your study.
  • Redacts transcripts turn by turn with timestamps and speaker labels preserved. Survey exports get new redacted columns; the originals are untouched.
  • Writes a report of what was replaced, by category and by layer, so a reviewer can audit it rather than assume.

What you get

Inputs
A transcript .txt from the transcription service, or a survey .csv export.
Output
A redacted copy plus the redaction report.
Schema
The 11-category scheme from ISR-BRIGHT, aligned to DHS 047-01-007 and U-M data sensitivity classifications.
Model
An open-weight language model served on the cluster. No text is sent to any external API.

// A redacted transcript is what our summarization step accepts. Unredacted input is refused.

Interested? Reach out.

Tell us about your recordings or survey text and the data classification they fall under. We will confirm the fit, set your study up, and show you how to submit jobs. Book a consult, send an inquiry, or email the team.

Campus resources

Expertise across campus.

ISR and U-M host deep expertise well beyond our team. If your problem fits one of these groups better, we will point you toward them. Grouped by what they do.

Analysis & study design

Statistical Consulting at ISR

The next chapter of CSCAR: campus-wide statistical consulting, launched at ISR in partnership with Statistics and Biostatistics.

Consulting details →

Statistical Design Group

SRC's free study-design service. Bring a design problem before you collect anything; leave with a defensible approach.

Statistical Design Group →

PDHP

Population Dynamics and Health Program: proposal and methods consulting, workshops, mentoring, and small grants for population science.

pdhp.isr.umich.edu →
Data collection

Survey Research Operations

SRC's data-collection unit: study design through fieldwork at local to national scale, including biomarkers, focus groups, and text coding.

SRO services →

NIMLAS

Methodological consulting and training for longitudinal studies of aging populations. Consulting is a member benefit rather than a general campus service.

nimlas.isr.umich.edu →

PSC Data Services

Population Studies Center cores for data acquisition, management, and survey methodology guidance.

psc.isr.umich.edu →
Data, computing & AI resources

ICPSR

Providing data resources and training in statistics, data analysis, and quantitative research methods since 1962.

icpsr.umich.edu →

MIDAS · AI Sandbox

Hands-on sessions to explore AI tools for research (no experience required), plus faculty AI consultations and pilot funding.

midas.umich.edu →

ITS ARC

Advanced Research Computing: consulting on research computing, plus the Great Lakes HPC cluster, GPU allocations, research storage, and secure computing environments.

ARC consulting →

// Listings describe publicly available services.

Get help

Three ways to get unstuck.

Not ready to book a consult? Send an inquiry, email the team, or just drop in.

GOOGLE FORM EMBEDS HERE

Project inquiry

Tell us about your project in a short form: what you're studying, what data you have, what you wish were automated. Responses land in our shared intake sheet and we reply within 2 business days.

MCOMMUNITY GROUP PENDING

Email the team

Prefer email? Reach the whole team at once through our group address. One message, everyone sees it, the right person answers.

[group-name]@umich.edu

Drop-in help · our longest-running program

CoderSpaces

Weekly virtual hubs open to the whole U-M community: faculty, staff, and students. Bring code, cluster problems, statistics questions, or a method you're just starting to learn. Hosts from departments across campus, and everyone is welcome regardless of skill level.

Between sessions, join our CoderSpaces Slack space on U-M Slack: self-join the "U-M CoderSpaces" MCommunity group (how-to) and you'll sync into the workspace within 24 hours, then log in here. Active U-M faculty, staff, and students only.

Want to host a session? Email Alexis.

Sponsored by ISR · ARC · CPS · PSC · CSCAR · LSA Technology Services

Tuesdays · 9:30-11 a.m. ET · Zoom
Wednesdays · 1:30-3 p.m. ET · Zoom
Public calendar (umich.edu login)
Sign-in sheet when you attend