Entrez le mot de passe pour continuer
One paragraph: I have not come with one idea, I have come with four. They all sit inside your team's core research (activity, behaviour, multi-sensor fusion, older adults) and all leverage the state of the art in 2026–2027 (video foundation models, multi-modal LLMs, self-supervised and continual learning). I would love your view on which one fits STARS best today.
A multi-modal foundation model for non-intrusive monitoring of older adults at home — behaviour and physiology together.
2 ClinicalNot just detecting falls — predicting them, hours to seconds ahead, from subtle gait, posture and breathing shifts.
3 VLM · RAGA vision-language model grounded on medical guidelines that writes short clinical notes from home video and wearables.
4 ContinualA behaviour model that keeps learning as the patient changes across months and years — without forgetting.
A multi-modal foundation model for continuous, non-intrusive behaviour and physiology monitoring of older adults at home. Full deep dive below (problem, vision, research questions, four technical modules, roadmap, deliverables, risks).
Populations are ageing fast. Dementia and frailty reach millions of families every year. Today, clinicians only see a patient a few times a year — and by then, the decline is often already visible.
At the same time, in 2026, AI has changed: video foundation models, multi-modal LLMs, and self-supervised pre-training now make it possible to learn general behaviour from unlabeled data.
The gap is real: continuous, in-home, contact-free observation of older adults, turned into early clinical signals. Nobody is doing this end-to-end yet.
Les populations vieillissent vite. La démence et la fragilité touchent des millions de familles chaque année. Aujourd'hui, les cliniciens ne voient un patient que quelques fois par an — et le déclin est souvent déjà visible à ce moment-là.
En même temps, en 2026, l'IA a changé : les modèles de fondation vidéo, les LLM multi-modaux et le pré-entraînement auto-supervisé permettent maintenant d'apprendre le comportement général à partir de données non annotées.
Le vide est réel : une observation continue, à domicile, sans contact des personnes âgées, transformée en signaux cliniques précoces. Personne ne le fait de bout en bout aujourd'hui.
One system with one foundation model, trained on RGB, depth, and wearable signals. From a single depth camera and one wearable, it extracts activities of daily living, breathing, heart rate, gait, and agitation — continuously.
It learns each person's personal baseline, and raises a signal when the baseline shifts. A vision-language model then writes a short weekly report for the doctor: what changed, why, and what to look at.
Un système avec un modèle de fondation, entraîné sur RGB, profondeur et signaux portables. À partir d'une seule caméra de profondeur et d'un capteur portable, il extrait les activités de la vie quotidienne, la respiration, la fréquence cardiaque, la marche et l'agitation — en continu.
Il apprend la référence personnelle de chaque personne, et lève un signal quand cette référence bouge. Un modèle vision-langage écrit ensuite un court rapport hebdomadaire pour le médecin : ce qui a changé, pourquoi, et sur quoi regarder.
Self-supervised pre-training on RGB + depth + IMU with masked auto-encoding and cross-modal contrastive learning. Backbone: video transformer à la InternVideo2 / VideoMAE-v2, extended with a depth stream and a wearable stream.
Breathing rate and tidal volume (from my PhD), heart rate via rPPG on face and hands, gait speed and posture stability — all decoded from the same depth camera as auxiliary heads.
Fine-tune the foundation model on Toyota Smarthome for activities of daily living, and on CoBTeK data for agitation, apathy, and wandering in dementia. Reuse STARS's skeleton and tracking pipelines.
A per-person self-supervised anomaly model over long windows. When the person's routine, gait, sleep, or breathing shift beyond their own baseline, we raise a graded clinical signal.
Bonus module — M5 — A vision-language reporter: a multi-modal LLM (LLaVA-style) that reads a week of signals and video snippets and writes a short structured note for the doctor, with evidence links back to the raw data.
Module bonus — M5 — Un rapporteur vision-langage : un LLM multi-modal (type LLaVA) qui lit une semaine de signaux et d'extraits vidéo et écrit une courte note structurée pour le médecin, avec des liens de preuve vers les données brutes.
Year 1 — 2026 — Foundations.
Year 2 — 2027 — Clinical validation and impact.
Année 1 — 2026 — Fondations.
Année 2 — 2027 — Validation clinique et impact.
STARS already leads on video understanding, activity recognition, and multi-sensor fusion — this is exactly the base I need to build on.
CoBTeK + Nice University Hospital gives the clinical data, the older-adult cohort, and the medical expertise. Without them, the project is just an academic paper.
3IA Côte d'Azur gives the AI infrastructure, compute, and the strategic frame. This work fits directly under the chair's health axis.
What I add: a working pipeline for contactless physiology from depth, hospital experience, a US patent I can extend, and hands-on bimodal deep learning from my recent fall-detection work.
STARS est déjà en pointe sur la compréhension vidéo, la reconnaissance d'activités et la fusion multi-capteurs — c'est exactement la base sur laquelle je veux construire.
CoBTeK + CHU de Nice apporte les données cliniques, la cohorte de personnes âgées et l'expertise médicale. Sans eux, le projet n'est qu'un papier académique.
3IA Côte d'Azur apporte l'infrastructure IA, le calcul et le cadre stratégique. Ce travail se place directement sous l'axe santé de la chaire.
Ce que j'ajoute : un pipeline opérationnel de physiologie sans contact par profondeur, une expérience hospitalière, un brevet américain que je peux étendre, et une pratique concrète du deep learning bimodal issue de mon travail récent sur la détection de chute.
“I want to build, with STARS and CoBTeK, a multi-modal foundation model that turns one depth camera and one wearable into an early-warning system for older adults — and lets a vision-language model explain, every week, what changed in a person's life.”
« Je veux construire, avec STARS et CoBTeK, un modèle de fondation multi-modal qui transforme une seule caméra de profondeur et un capteur portable en système d'alerte précoce pour les personnes âgées — et qui laisse un modèle vision-langage expliquer, chaque semaine, ce qui a changé dans la vie d'une personne. »
Falls are the number-one cause of hospitalisation in older adults. Everyone builds fall detection; almost nobody builds fall prediction. This project turns my recent bimodal work into a proactive alerting system.
Vision. Predict a fall hours before from subtle daily changes (gait speed, stride variability, breathing rhythm, sit-to-stand time), and seconds before from micro-instabilities in posture.
Why STARS. This is a direct extension of my upcoming IEEE ICM 2026 paper on camera + IMU fusion, and it sits inside STARS's core work on human motion. CoBTeK gives access to older-adult data.
Approach:
Roadmap. Y1 (2026): dataset extension, prediction model, first paper. Y2 (2027): clinical pilot at CoBTeK, wearable alerting demo, second paper + grant application.
Deliverables. 2 papers, open dataset extension, real-time demo, patent option on the prediction pipeline.
Vision. Prédire une chute des heures avant à partir de changements subtils du quotidien (vitesse de marche, variabilité du pas, rythme respiratoire, temps assis-debout), et des secondes avant à partir de micro-instabilités de la posture.
Pourquoi STARS. C'est une extension directe de mon papier IEEE ICM 2026 sur la fusion caméra + IMU, et cela s'inscrit dans le cœur du travail de STARS sur le mouvement humain. CoBTeK donne accès aux données de personnes âgées.
Approche :
Feuille de route. A1 (2026) : extension du jeu de données, modèle de prédiction, premier papier. A2 (2027) : pilote clinique à CoBTeK, démo d'alerte portable, deuxième papier + demande de financement.
Livrables. 2 papiers, extension de jeu de données ouvert, démo temps réel, option de brevet sur le pipeline de prédiction.
Nurses spend one third of their time writing notes. Family caregivers rarely track anything reliably. This project trains a healthcare-specific vision-language model that watches, understands, and writes.
Vision. A vision-language model that ingests a day of home video and wearable signals and produces a short, structured clinical note: "08:15 — medication taken, gait unsteady, mild agitation observed", with links back to the raw evidence.
Why STARS. STARS already builds skeleton and activity pipelines — the perfect visual backbone. Adding an LLM head with RAG on medical guidelines (SNOMED, ICD-10, geriatric care protocols) gives clinicians something they can trust.
Approach:
Roadmap. Y1 (2026): build the model, evaluate on Toyota Smarthome + CoBTeK. Y2 (2027): clinician-in-the-loop trial, publish + demo at a nursing home.
Deliverables. 2 papers (one methods, one clinical), open medical VLM benchmark, prototype dashboard for nursing homes.
Vision. Un modèle vision-langage qui absorbe une journée de vidéo à domicile et de signaux portables et produit une note clinique courte et structurée : « 08h15 — médicament pris, marche instable, agitation légère observée », avec des liens vers les preuves brutes.
Pourquoi STARS. STARS construit déjà les pipelines squelette et activité — l'ossature visuelle parfaite. Ajouter une tête LLM avec RAG sur les guidelines médicaux (SNOMED, ICD-10, protocoles gériatriques) donne aux cliniciens quelque chose de fiable.
Approche :
Feuille de route. A1 (2026) : construction du modèle, évaluation sur Toyota Smarthome + CoBTeK. A2 (2027) : essai clinicien-en-boucle, publication + démo en EHPAD.
Livrables. 2 papiers (méthodes + clinique), benchmark VLM médical ouvert, tableau de bord prototype pour EHPAD.
Brémond lists life-long learning explicitly in his research topics. Static models fail on ageing patients because their routine, gait and cognition drift for real. This project makes the model drift with them — without forgetting.
Vision. A behaviour model that continuously adapts to each older adult over months and years — and detects meaningful drift (early cognitive decline, disease progression) versus normal ageing.
Why STARS. This maps directly onto Brémond's stated interest in life-long learning, spatio-temporal reasoning and uncertainty. CoBTeK's longitudinal cohort is the ideal testbed.
Approach:
Roadmap. Y1 (2026): representation and continual learning, first paper. Y2 (2027): 12-month CoBTeK study, drift-based decline prediction, second paper.
Deliverables. 2 papers, longitudinal benchmark protocol, open code, clinical drift dashboard.
Vision. Un modèle de comportement qui s'adapte en continu à chaque personne âgée sur des mois et des années — et détecte les dérives significatives (déclin cognitif précoce, progression de maladie) par rapport au vieillissement normal.
Pourquoi STARS. Cela cible directement l'intérêt déclaré de Brémond pour le life-long learning, le raisonnement spatio-temporel et la gestion de l'incertitude. La cohorte longitudinale de CoBTeK est le terrain de test idéal.
Approche :
Feuille de route. A1 (2026) : représentation et apprentissage continu, premier papier. A2 (2027) : étude CoBTeK 12 mois, prédiction de déclin par dérive, deuxième papier.
Livrables. 2 papiers, protocole de benchmark longitudinal, code ouvert, tableau de bord clinique de dérive.
My honest answer: Idea 1 — BehaviourWatch-3D — is the flagship. It absorbs the other three as sub-modules and generates the highest number of top-tier papers.
But if you tell me you need a fast, focused win first, I would start with Idea 2 — Fall-Ahead-3D: 12 months to a first paper and a clinical demo, directly building on my IEEE ICM 2026 submission.
If you tell me you want to plant a flag in a hot 2026 space, we go with Idea 3 — MedActionLLM: it puts STARS on the map of medical vision-language models.
If you tell me the priority is what you already care about most, we go with Idea 4 — LifelongCare: it lands squarely on your stated interest in life-long learning.
I am ready to lead any of them — and honestly, I would love to hear which one excites you.
Ma réponse honnête : l'idée 1 — BehaviourWatch-3D — est le projet phare. Elle absorbe les trois autres comme sous-modules et produit le plus grand nombre de papiers de haut niveau.
Mais si vous me dites qu'il faut une victoire rapide et ciblée d'abord, je commence par l'idée 2 — Fall-Ahead-3D : 12 mois pour un premier papier et une démo clinique, directement dans la continuité de ma soumission IEEE ICM 2026.
Si vous me dites qu'il faut planter un drapeau dans un espace chaud en 2026, on part sur l'idée 3 — MedActionLLM : elle place STARS sur la carte des modèles vision-langage médicaux.
Si vous me dites que la priorité est ce qui compte déjà le plus pour vous, on part sur l'idée 4 — LifelongCare : elle atterrit pile sur votre intérêt déclaré pour le life-long learning.
Je suis prêt à mener n'importe laquelle — et honnêtement, j'aimerais entendre celle qui vous excite le plus.