Обучение с подкреплением (ИИ24, 7 модуль): различия между версиями

Материал из Wiki - Факультет компьютерных наук
Перейти к навигации Перейти к поиску
Ссылки на записи
Ссылки
Строка 29: Строка 29:
| style="background:#eaecf0;" | '''1''' [[https://www.youtube.com/watch?v=KLEcPmdR87U&list=PLmA-1xX7IuzDe8CEWijYwsgmdHXyaEQsg&index=5&pp=iAQB Youtube]] [[https://vkvideo.ru/playlist/-227011779_68/video-227011779_456239732?linked=1 VKVideo]] || Intro to RL, Dynamic Programming  || 10/01/2026 ||  
| style="background:#eaecf0;" | '''1''' [[https://www.youtube.com/watch?v=KLEcPmdR87U&list=PLmA-1xX7IuzDe8CEWijYwsgmdHXyaEQsg&index=5&pp=iAQB Youtube]] [[https://vkvideo.ru/playlist/-227011779_68/video-227011779_456239732?linked=1 VKVideo]] || Intro to RL, Dynamic Programming  || 10/01/2026 ||  
|-
|-
| style="background:#eaecf0;" | '''2''' [[ Запись]] || Model-Free Tabular RL: Q-Learning, SARSA || 17/01/2026 ||  
| style="background:#eaecf0;" | '''2''' [[https://www.youtube.com/watch?v=Uf-KHdRh3zs&list=PLmA-1xX7IuzDe8CEWijYwsgmdHXyaEQsg&index=4&pp=iAQB Youtube]] [[https://vkvideo.ru/playlist/-227011779_68/video-227011779_456239762?linked=1 VKVideo]] || Model-Free Tabular RL: Q-Learning, SARSA || 17/01/2026 ||  
|-
|-
| style="background:#eaecf0;" | '''3''' [[ Запись]] || Intro to Deep RL: DQN, RAINBOW and beyond || 24/01/2026 ||  
| style="background:#eaecf0;" | '''3''' [[ Youtube]] [[ VKVideo]] || Intro to Deep RL: DQN, RAINBOW and beyond || 24/01/2026 ||  
|-
|-
| style="background:#eaecf0;" | '''4''' [[ Запись]] || Policy-Based Methods: Policy Gradient, REINFORCE, A2C || 31/01/2026 ||  
| style="background:#eaecf0;" | '''4''' [[ Youtube]] [[ VKVideo]] || Policy-Based Methods: Policy Gradient, REINFORCE, A2C || 31/01/2026 ||  
|-
|-
| style="background:#eaecf0;" | '''5''' [[ Запись]] || Advanced Policy-Based: TRPO, PPO and beyond || 07/02/2026 ||  
| style="background:#eaecf0;" | '''5''' [[ Youtube]] [[ VKVideo]] || Advanced Policy-Based: TRPO, PPO and beyond || 07/02/2026 ||  
|-
|-
| style="background:#eaecf0;" | '''6''' [[ Запись]] || Continuous Control: DDPG, SAC and beyond || 14/02/2026 ||  
| style="background:#eaecf0;" | '''6''' [[ Youtube]] [[ VKVideo]] || Continuous Control: DDPG, SAC and beyond || 14/02/2026 ||  
|-
|-
| style="background:#eaecf0;" | '''7''' [[ Запись]] || Offline RL || 21/02/2026 ||  
| style="background:#eaecf0;" | '''7''' [[ Youtube]] [[ VKVideo]] || Offline RL || 21/02/2026 ||  
|-
|-
| style="background:#eaecf0;" | '''8''' [[ Запись]] || Multi-Armed Bandits || 28/02/2026 ||  
| style="background:#eaecf0;" | '''8''' [[ Youtube]] [[ VKVideo]] || Multi-Armed Bandits || 28/02/2026 ||  
|-
|-
| style="background:#eaecf0;" | '''9''' [[ Запись]] || Model-based RL: AlphaZero and friends || 07/03/2026 ||  
| style="background:#eaecf0;" | '''9''' [[ Youtube]] [[ VKVideo]] || Model-based RL: AlphaZero and friends || 07/03/2026 ||  
|-
|-
| style="background:#eaecf0;" | '''10''' [[ Запись]] || RL in a context of LLMs || 14/03/2026 ||  
| style="background:#eaecf0;" | '''10''' [[ Youtube]] [[ VKVideo]] || RL in a context of LLMs || 14/03/2026 ||  
|-
|-
| style="background:#eaecf0;" | '''11''' [[ Запись]] || Practical RL || 21/03/2026 ||  
| style="background:#eaecf0;" | '''11''' [[ Youtube]] [[ VKVideo]] || Practical RL || 21/03/2026 ||  
|-
|-
|}
|}
=== Записи консультаций ===


==Формула оценивания==
==Формула оценивания==

Версия от 12:36, 14 февраля 2026

О курсе

Занятия проводятся в Zoom по субботам с 13:00 МСК

Контакты

Чат курса в TG: [[1]]

Преподаватель: Сергей Лактионов, Вячеслав Бучков

Ассистент Контакты

Материалы курса

Ссылка на плейлист курса на YouTube: [YouTube-playlist]

Ссылка на GitHub с материалами курса: [GitHub repository]

Занятие Тема Дата Дополнительные материалы
1 [Youtube] [VKVideo] Intro to RL, Dynamic Programming 10/01/2026
2 [Youtube] [VKVideo] Model-Free Tabular RL: Q-Learning, SARSA 17/01/2026
3 Youtube VKVideo Intro to Deep RL: DQN, RAINBOW and beyond 24/01/2026
4 Youtube VKVideo Policy-Based Methods: Policy Gradient, REINFORCE, A2C 31/01/2026
5 Youtube VKVideo Advanced Policy-Based: TRPO, PPO and beyond 07/02/2026
6 Youtube VKVideo Continuous Control: DDPG, SAC and beyond 14/02/2026
7 Youtube VKVideo Offline RL 21/02/2026
8 Youtube VKVideo Multi-Armed Bandits 28/02/2026
9 Youtube VKVideo Model-based RL: AlphaZero and friends 07/03/2026
10 Youtube VKVideo RL in a context of LLMs 14/03/2026
11 Youtube VKVideo Practical RL 21/03/2026

Формула оценивания

Оценка = МИН(10, 10*(0.65*HW + 0.10*TA + 0.25*RC)), где

  • HW - сумма баллов за (как минимум) 5 ДЗ;
  • RC - оценка за презентацию статьи, посвященной новым алгоритмам или неожиданным применениям RL-парадигмы в индустрии;
  • TA - сумма баллов за еженедельные квизы (суммарно 10 квизов).

Для каждого домашнего задания есть мягкий дедлайн, сдача после которого в течение недели до жёсткого дедлайна оценивается со штрафом 5% от оценки за ДЗ за каждый день просрочки.

Домашние задания

Литература