Skip to content

Accumulated changes to the integration branch - #1596

Merged
qurat08 merged 545 commits into
mainfrom
dev
Sep 30, 2026
Merged

qurat08 merged 545 commits into
mainfrom
dev

Conversation

@ralf-berger

Copy link
Copy Markdown
Collaborator

Branching strategy: Everything gets merged into dev first, then merged again into main.

dependabot Bot and others added 30 commits February 18, 2025 10:00
Bumps nginxinc/nginx-unprivileged from 1.27.3-alpine to 1.27.4-alpine.

---
updated-dependencies:
- dependency-name: nginxinc/nginx-unprivileged
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
…MKG"

courseId was undefined in the HTTP request.
Prepared the logging for the activity "User filtered top N concepts".
It doesn't work yet.
…w/ understood/ not understood.

-Fixed the problem of the activity logging of "User viewed recommended concepts"
-edited the type of the recommended concepts. Previously they could be either main_concept/ related_concept or both.
-Now the ones that are under recommended concepts Tab, they are considered as recommended_concept.
-Added type to the concepts while marking as understood/ not understood/ new.
-There is still a logging problem, by marking a concept from the recommended materials tab!
…mended concept"

-Additionally, commented out some unnecessary console logs and added comments where needed.
Two activites are being logged by adding an annotation:
"User added an annotation" and "User annotated a material".
The annotation object can be "Note", "Question", "External Resource".
The material object can be "pdf", "video", "Youtube".
So the 6 possibilities are the following:
"User added a note",
"User added an external resource",
"User asked a question",
"User annotated a PDF",
"User annotated a video",
"User annotated a youtube video".
The activities include "User zoomed in a pdf", "User zoomed out a pdf", "User reset zoom in a pdf"
-Added comments for the previous implementation
-Removed console logs.
The activity is "User viewed a Material's slide"
-Logged the Activity "User did not understand a slide", additionally to "User accessed Slide Kg"
Activities include: "User marked a notification as read",  "User marked a notification as unread"
fixing a typo and adding some text
Activities include:
"User viewed notifications "
"User marked a notification/s as read",
"User marked a notification/s as unread",
"User starred a notification/s",
"User unstarred a notification/s",
"User deleted a notification/s"
Activities include:
"User follow/ unfollow/ hid/ unhid/ filtered/ replied/ added/asked"
Implementation include the following activities:
"User accessed Slide Knowledge Graph",
"User accessed Material Knowledge Graph",
"User accessed Course Knowledge Graph",
"User did not understand a Slide".
The activities include:
"User uploaded/ annotated/ zoomed in/ zoomed out/ reset zoom a pdf".
"User annotated/ uploaded a video".
"User annotated/ uploaded a youtube video".
Bumps [axios](https://github.com/axios/axios) from 1.7.9 to 1.8.0.
- [Release notes](https://github.com/axios/axios/releases)
- [Changelog](https://github.com/axios/axios/blob/v1.x/CHANGELOG.md)
- [Commits](axios/axios@v1.7.9...v1.8.0)

---
updated-dependencies:
- dependency-name: axios
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Bumps [mongoose](https://github.com/Automattic/mongoose) from 8.10.1 to 8.10.2.
- [Release notes](https://github.com/Automattic/mongoose/releases)
- [Changelog](https://github.com/Automattic/mongoose/blob/master/CHANGELOG.md)
- [Commits](Automattic/mongoose@8.10.1...8.10.2)

---
updated-dependencies:
- dependency-name: mongoose
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 10

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (2)
coursemapper-kg/recommendation/app/services/course_materials/recommendation/youtube_service.py (2)

33-35: Do not set GOOGLE_APPLICATION_CREDENTIALS in code

Remove inline env mutation; configure via deployment (env/secret manager). Leaving it here risks leaking paths and breaks local/runtime separation.

-    os.environ[
-        "GOOGLE_APPLICATION_CREDENTIALS"
-    ] = "masterthesis-350015-47ab14d0b53b.json"
+    # GOOGLE_APPLICATION_CREDENTIALS must be provided by the environment (CI/K8s/secret store).

176-178: Avoid global string fill; fill per-column with typed defaults

Global "-1" converts numerics to strings.

-            video_data = video_data.fillna("-1")
+            video_data = video_data.fillna({
+                "duration": "0",
+                "views": 0,
+                "description_full": "",
+                "like_count": 0,
+                "channel_title": "",
+                "text": "",
+            })
♻️ Duplicate comments (3)
coursemapper-kg/recommendation/app/services/course_materials/recommendation/youtube_service.py (3)

39-52: Hard-coded API keys committed — rotate immediately and load from env/secret store

All keys must be revoked and removed from source. Load a comma-separated list from env and fail fast if missing. Also drop the unused class-level client to avoid accidental usage.

-    # DEVELOPER_KEY = os.environ.get("YOUTUBE_API_KEY")
-    DEVELOPER_KEY = "AIzaSyBphZOn7EJmPMmZwrB71aepaA5Rbuex9MU"
-    youtube = googleapiclient.discovery.build(
-        api_service_name,
-        api_version,
-        developerKey="AIzaSyClxnNwQ1x34pGioQazLlGxOjO9Fp2GGTY",
-    )
-    DEVELOPER_KEYS = [
-        "AIzaSyD_CGmR_Voq4DIV5okRaR6G8adoe-ZSZsM",
-        "AIzaSyClxnNwQ1x34pGioQazLlGxOjO9Fp2GGTY",
-        "AIzaSyADNntK6m7DbA6eZFYOa9Y8e6IYHykUUFE",
-        "AIzaSyBphZOn7EJmPMmZwrB71aepaA5Rbuex9MU",
-        "AIzaSyB2Wck31LUlgsqI7dgTcC2dMeeVXgb9TDI",
-    ]
+    # Comma-separated API keys from env (e.g., "k1,k2,k3"). Do NOT commit keys.
+    DEVELOPER_KEYS = [k.strip() for k in os.getenv("YOUTUBE_API_KEYS", "").split(",") if k.strip()]
+    if not DEVELOPER_KEYS:
+        raise RuntimeError("YOUTUBE_API_KEYS env var is required (comma-separated YouTube API keys).")

54-90: Retry/backoff logic is broken; fix loop, logging, and quota handling

‘i’ never increments; retries never happen; retry_count == 0 is unreachable; use logger.exception and switch keys on quotaExceeded.

-    def search_youtube_videos(self, developer_keys, query, top_n=50, api_service_name="youtube", api_version="v3"):
+    def search_youtube_videos(self, developer_keys, query, top_n=50, api_service_name="youtube", api_version="v3"):
         """
             Switching YouTube API keys
         """
-        retry_count = 3
-        retry_delay = 5
-        i = 0
-        for key in developer_keys:
-            try:
-                youtube = googleapiclient.discovery.build(api_service_name, api_version, developerKey=key)
-                request = youtube.search().list(
-                    part="snippet",
-                    maxResults=top_n,
-                    type="video",
-                    q=query,
-                    relevanceLanguage="en",
-                )
-                return request.execute(), youtube
-            except (ConnectionAbortedError, ConnectionResetError, timeout) as e:
-                logger.error("Error while getting the videos")
-                logger.error(e)
-                if i == retry_count - 1:
-                    raise  # re-raise the exception if all retries fail
-                delay = retry_delay * (2 ** i)  # use a backoff algorithm to increase the delay
-                time.sleep(delay)
-                logger.info("New Try")
-                if retry_count == 0:
-                    return None, None
-            except HttpError as e:
-                if e.resp.status == 403 and "quota" in str(e):
-                    print(f"Quota exceeded for key: {key}. Trying next key...")
-                else:
-                    raise e
-        raise Exception("All API keys have exceeded their quota.")
+        retry_count = 3
+        retry_delay = 5
+        top_n = min(int(top_n), 50)  # API max
+        developer_keys = developer_keys or self.DEVELOPER_KEYS
+        for key in developer_keys:
+            for i in range(retry_count):
+                try:
+                    youtube = googleapiclient.discovery.build(api_service_name, api_version, developerKey=key)
+                    request = youtube.search().list(
+                        part="snippet",
+                        maxResults=top_n,
+                        type="video",
+                        q=query,
+                        relevanceLanguage="en",
+                    )
+                    return request.execute(), youtube
+                except (ConnectionAbortedError, ConnectionResetError, socket.timeout) as e:
+                    logger.exception("Transient connection error on attempt %d/%d (key ****%s)", i + 1, retry_count, key[-4:])
+                    time.sleep(retry_delay * (2 ** i))
+                    continue
+                except HttpError as e:
+                    status = getattr(getattr(e, "resp", None), "status", None)
+                    msg = str(e).lower()
+                    if status == 403 and ("quotaexceeded" in msg or "dailylimitexceeded" in msg or "quota" in msg):
+                        logger.warning("Quota exceeded for key ****%s; switching key…", key[-4:])
+                        break  # next key
+                    raise
+        raise RuntimeError("All YouTube API keys exhausted or failed.")

Additionally apply outside this range:

# at top of file
import socket  # replace 'from socket import *'

# and remove the star import usage in excepts (done in diff above).

73-76: Log exceptions with traceback; avoid double error lines

Use logger.exception once; also avoid relying on timeout from star-import.

-            except (ConnectionAbortedError, ConnectionResetError, timeout) as e:
-                logger.error("Error while getting the videos")
-                logger.error(e)
+            except (ConnectionAbortedError, ConnectionResetError, socket.timeout) as e:
+                logger.exception("Error while getting the videos")

Outside this range, replace from socket import * with import socket.

📜 Review details

Configuration used: CodeRabbit UI

Review profile: ASSERTIVE

Plan: Pro

💡 Knowledge Base configuration:

  • MCP integration is disabled by default for public repositories
  • Jira integration is disabled by default for public repositories
  • Linear integration is disabled by default for public repositories

You can enable these sources in your CodeRabbit configuration.

📥 Commits

Reviewing files that changed from the base of the PR and between 63591bd and 4b97a54.

📒 Files selected for processing (1)
  • coursemapper-kg/recommendation/app/services/course_materials/recommendation/youtube_service.py (3 hunks)
🧰 Additional context used
🪛 Gitleaks (8.27.2)
coursemapper-kg/recommendation/app/services/course_materials/recommendation/youtube_service.py

47-47: Uncovered a GCP API key, which could lead to unauthorized access to Google Cloud services and data breaches.

(gcp-api-key)


48-48: Uncovered a GCP API key, which could lead to unauthorized access to Google Cloud services and data breaches.

(gcp-api-key)


49-49: Uncovered a GCP API key, which could lead to unauthorized access to Google Cloud services and data breaches.

(gcp-api-key)


50-50: Uncovered a GCP API key, which could lead to unauthorized access to Google Cloud services and data breaches.

(gcp-api-key)


51-51: Uncovered a GCP API key, which could lead to unauthorized access to Google Cloud services and data breaches.

(gcp-api-key)

⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (4)
  • GitHub Check: Build container images / Build image (coursemapper-kg/preprocess)
  • GitHub Check: Build container images / Build image (coursemapper-kg/concept-map)
  • GitHub Check: Build container images / Build image (coursemapper-kg/recommendation)
  • GitHub Check: Build container images / Build image (webapp)
🔇 Additional comments (1)
coursemapper-kg/recommendation/app/services/course_materials/recommendation/youtube_service.py (1)

193-199: Good: exception visibility improved

Catching HttpError and logging with traceback is appropriate.

import logging
from log import LOG
import time
from googleapiclient.errors import HttpError

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick (assertive)

Minor: clean imports

Now that HttpError is imported directly, drop import googleapiclient.errors. Also remove duplicate import time (present twice).

-from youtube_transcript_api import YouTubeTranscriptApi
+from youtube_transcript_api import YouTubeTranscriptApi
@@
-import googleapiclient.errors
+# googleapiclient.errors not needed; using HttpError directly
@@
-import time
+# remove duplicate 'import time' if present twice

Committable suggestion skipped: line range outside the PR's diff.

🤖 Prompt for AI Agents
In
coursemapper-kg/recommendation/app/services/course_materials/recommendation/youtube_service.py
around line 15, tidy imports by removing the redundant module import and
duplicate import: delete the now-unnecessary "import googleapiclient.errors"
since HttpError is imported directly from googleapiclient.errors, and remove the
duplicated "import time" so time is only imported once; ensure no other
references depend on the removed module import.

Comment on lines +84 to +88
except HttpError as e:
if e.resp.status == 403 and "quota" in str(e):
print(f"Quota exceeded for key: {key}. Trying next key...")
else:
raise e

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🛠️ Refactor suggestion

Replace print with structured logging; broaden quota reasons

Use logger and check standard reasons.

-                if e.resp.status == 403 and "quota" in str(e):
-                    print(f"Quota exceeded for key: {key}. Trying next key...")
+                if e.resp.status == 403 and any(r in str(e).lower() for r in ["quotaexceeded", "dailylimitexceeded", "quota"]):
+                    logger.warning("Quota exceeded for key ****%s. Trying next key…", key[-4:])
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
except HttpError as e:
if e.resp.status == 403 and "quota" in str(e):
print(f"Quota exceeded for key: {key}. Trying next key...")
else:
raise e
except HttpError as e:
if e.resp.status == 403 and any(r in str(e).lower() for r in ["quotaexceeded", "dailylimitexceeded", "quota"]):
logger.warning("Quota exceeded for key ****%s. Trying next key…", key[-4:])
else:
raise e
🤖 Prompt for AI Agents
In
coursemapper-kg/recommendation/app/services/course_materials/recommendation/youtube_service.py
around lines 84-88, replace the plain print with structured logging and broaden
the quota detection: use the module or instance logger (e.g., logger.warning) to
log a clear message that includes the key and exception details (include
exc_info or the exception object), and detect quota errors not only by checking
if "quota" is in the string but also by inspecting standard API error reasons
such as "quotaExceeded", "dailyLimitExceeded" or the HttpError response
body/reason; if a quota-related condition is detected, log the warning and
continue to the next key, otherwise re-raise the exception.

Comment on lines +104 to +106
response, youtube_api_sinlge = self.search_youtube_videos(
developer_keys=self.DEVELOPER_KEYS, query=concepts, top_n=top_n
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick (assertive)

Typo: youtube_api_sinlge → youtube_api_single

Fix naming for readability and to avoid propagating typos.

-        response, youtube_api_sinlge = self.search_youtube_videos(
+        response, youtube_api_single = self.search_youtube_videos(
             developer_keys=self.DEVELOPER_KEYS, query=concepts, top_n=top_n
         )
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
response, youtube_api_sinlge = self.search_youtube_videos(
developer_keys=self.DEVELOPER_KEYS, query=concepts, top_n=top_n
)
response, youtube_api_single = self.search_youtube_videos(
developer_keys=self.DEVELOPER_KEYS, query=concepts, top_n=top_n
)
🤖 Prompt for AI Agents
In
coursemapper-kg/recommendation/app/services/course_materials/recommendation/youtube_service.py
around lines 104 to 106, there's a typo in the variable name
"youtube_api_sinlge" — rename it to "youtube_api_single" consistently where it's
assigned and anywhere else it's referenced to improve readability and avoid
further typos; update the function return unpacking to use youtube_api_single
and search for other occurrences of the misspelled identifier in the file and
replace them to keep names consistent.

Comment on lines 108 to 111
if len(response["items"]) == 0:
logger.info("No Video found for this input")
return []
else:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🛠️ Refactor suggestion

Guard against None/empty responses and return consistent type

Return an empty DataFrame (not list) to keep a stable return type.

-        if len(response["items"]) == 0:
+        if not response or not response.get("items"):
             logger.info("No Video found for this input")
-            return []
+            return pd.DataFrame()
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
if len(response["items"]) == 0:
logger.info("No Video found for this input")
return []
else:
if not response or not response.get("items"):
logger.info("No Video found for this input")
return pd.DataFrame()
else:
🤖 Prompt for AI Agents
In
coursemapper-kg/recommendation/app/services/course_materials/recommendation/youtube_service.py
around lines 108 to 111, the code returns an empty list when no videos are
found; change this to return an empty pandas DataFrame to maintain a consistent
return type. Guard against response being None or missing the "items" key before
accessing it (e.g., check response and "items" in response), and return
pandas.DataFrame() when there are no items or response is invalid so callers
always receive a DataFrame.

Comment on lines +95 to +137
for index, id in enumerate(df_ids["id"]):
# try:
# df_snippet["text"][index] = df_snippet["text"][index] + ". " + get_subtitles(id)
# df_snippet["text"][index] = df_snippet["text"][index] + ". " + get_subtitles(id)
# except (NoTranscriptFound, TranscriptsDisabled) as e:
# logger.error("No transcript found in english or transcript disabled for this video "
# "https://www.youtube.com/watch?v={} ".format(id))
# logger.error("No transcript found in english or transcript disabled for this video "
# "https://www.youtube.com/watch?v={} ".format(id))

try:
duration, views, description = self.get_video_details(id)
duration = re.findall(r"\d+", duration)
duration = ":".join(duration)
# print(id, duration, views)
duration_list.append(duration)
view_list.append(views)
description_list.append(description)
except Exception as e:
logger.error("Error while getting the videos details", e)
res = self.get_video_details(youtube_api_sinlge, id)
if res is None:
raise ValueError("No details returned for video id {}".format(id))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick (assertive)

Avoid shadowing built-in id and improve error

Rename loop var to video_id; keep error message intact.

-            for index, id in enumerate(df_ids["id"]):
+            for index, video_id in enumerate(df_ids["id"]):
@@
-                    res = self.get_video_details(youtube_api_sinlge, id)
+                    res = self.get_video_details(youtube_api_single, video_id)
-                    if res is None:
-                        raise ValueError("No details returned for video id {}".format(id))
+                    if res is None:
+                        raise ValueError("No details returned for video id {}".format(video_id))
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
for index, id in enumerate(df_ids["id"]):
# try:
# df_snippet["text"][index] = df_snippet["text"][index] + ". " + get_subtitles(id)
# df_snippet["text"][index] = df_snippet["text"][index] + ". " + get_subtitles(id)
# except (NoTranscriptFound, TranscriptsDisabled) as e:
# logger.error("No transcript found in english or transcript disabled for this video "
# "https://www.youtube.com/watch?v={} ".format(id))
# logger.error("No transcript found in english or transcript disabled for this video "
# "https://www.youtube.com/watch?v={} ".format(id))
try:
duration, views, description = self.get_video_details(id)
duration = re.findall(r"\d+", duration)
duration = ":".join(duration)
# print(id, duration, views)
duration_list.append(duration)
view_list.append(views)
description_list.append(description)
except Exception as e:
logger.error("Error while getting the videos details", e)
res = self.get_video_details(youtube_api_sinlge, id)
if res is None:
raise ValueError("No details returned for video id {}".format(id))
for index, video_id in enumerate(df_ids["id"]):
# try:
# df_snippet["text"][index] = df_snippet["text"][index] + ". " + get_subtitles(id)
# except (NoTranscriptFound, TranscriptsDisabled) as e:
# logger.error("No transcript found in english or transcript disabled for this video "
# "https://www.youtube.com/watch?v={} ".format(id))
try:
- res = self.get_video_details(youtube_api_sinlge, id)
- if res is None:
res = self.get_video_details(youtube_api_single, video_id)
if res is None:
raise ValueError("No details returned for video id {}".format(video_id))
🤖 Prompt for AI Agents
In
coursemapper-kg/recommendation/app/services/course_materials/recommendation/youtube_service.py
around lines 127 to 137, the loop variable currently named "id" shadows Python's
built-in id() and should be renamed to "video_id"; update the for statement and
every reference inside the loop (e.g., pass video_id to get_video_details and to
the error message string.format) so behavior and the existing error message
remain the same while avoiding built-in shadowing.

Comment on lines +141 to +153
# Keep your duration normalization
duration = re.findall(r"\d+", str(duration))
duration = ":".join(duration)
except Exception:
# Fix logging formatter error and keep same message semantics
logger.exception("Error while getting the videos details for id %s", id)
# Append safe defaults to keep list lengths aligned with df_ids
duration = "0"
views = 0
description = ""
like_count = 0
channel_title = ""

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick (assertive)

Normalize duration properly; keep numeric types for counts

Parse ISO-8601 with isodate if available; coerce counts to ints.

-                    duration = re.findall(r"\d+", str(duration))
-                    duration = ":".join(duration)
+                    # Prefer true ISO-8601 parsing if available; fallback kept
+                    try:
+                        import isodate
+                        duration = int(isodate.parse_duration(str(duration)).total_seconds())
+                    except Exception:
+                        duration = ":".join(re.findall(r"\d+", str(duration)))
@@
-                    views = 0
+                    views = 0
@@
-                    like_count = 0
+                    like_count = 0

And when appending (below), coerce to numeric (see next comment).

📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
# Keep your duration normalization
duration = re.findall(r"\d+", str(duration))
duration = ":".join(duration)
except Exception:
# Fix logging formatter error and keep same message semantics
logger.exception("Error while getting the videos details for id %s", id)
# Append safe defaults to keep list lengths aligned with df_ids
duration = "0"
views = 0
description = ""
like_count = 0
channel_title = ""
# Keep your duration normalization
- duration = re.findall(r"\d+", str(duration))
# Prefer true ISO-8601 parsing if available; fallback to regex join
try:
import isodate
duration = int(isodate.parse_duration(str(duration)).total_seconds())
except Exception:
duration = ":".join(re.findall(r"\d+", str(duration)))
except Exception:
# Fix logging formatter error and keep same message semantics
logger.exception("Error while getting the videos details for id %s", id)
# Append safe defaults to keep list lengths aligned with df_ids
duration = "0"
views = 0
description = ""
like_count = 0
channel_title = ""
🤖 Prompt for AI Agents
In
coursemapper-kg/recommendation/app/services/course_materials/recommendation/youtube_service.py
around lines 141 to 153, the duration and numeric fields are being treated as
strings and not normalized consistently; update the exception handling and
normal parsing to (1) parse ISO-8601 durations using isodate.parse_duration
(fall back to existing regex only if isodate is unavailable) and normalize to a
canonical "H:MM:SS" or "MM:SS" string for duration, (2) ensure counts like views
and like_count are coerced to int (use int(...) with safe default 0) and
description/channel_title stay as strings, and (3) keep the existing
logger.exception call but maintain its safe formatting; after fixing here,
ensure any append/population logic below appends ints for numeric fields and the
normalized duration string so list lengths and types remain consistent.

Comment on lines +163 to +169
# Assign lists (same behavior as before, now lengths aligned)
video_data["duration"] = pd.Series(duration_list)
video_data["views"] = pd.Series(view_list)
video_data["description_full"] = pd.Series(description_list)
video_data["like_count"] = pd.Series(like_count_list)
video_data["channel_title"] = pd.Series(channel_title_list)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🛠️ Refactor suggestion

Preserve numeric dtypes for views/likes; avoid implicit object dtype

Cast with to_numeric and nullable Int64.

-            video_data["views"] = pd.Series(view_list)
+            video_data["views"] = pd.to_numeric(pd.Series(view_list), errors="coerce").fillna(0).astype("Int64")
@@
-            video_data["like_count"] = pd.Series(like_count_list)
+            video_data["like_count"] = pd.to_numeric(pd.Series(like_count_list), errors="coerce").fillna(0).astype("Int64")
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
# Assign lists (same behavior as before, now lengths aligned)
video_data["duration"] = pd.Series(duration_list)
video_data["views"] = pd.Series(view_list)
video_data["description_full"] = pd.Series(description_list)
video_data["like_count"] = pd.Series(like_count_list)
video_data["channel_title"] = pd.Series(channel_title_list)
# Assign lists (same behavior as before, now lengths aligned)
video_data["duration"] = pd.Series(duration_list)
video_data["views"] = pd.to_numeric(pd.Series(view_list), errors="coerce").fillna(0).astype("Int64")
video_data["description_full"] = pd.Series(description_list)
video_data["like_count"] = pd.to_numeric(pd.Series(like_count_list), errors="coerce").fillna(0).astype("Int64")
video_data["channel_title"] = pd.Series(channel_title_list)
🤖 Prompt for AI Agents
In
coursemapper-kg/recommendation/app/services/course_materials/recommendation/youtube_service.py
around lines 163 to 169, assigning views and like_count directly from lists
produces object dtype; convert these Series to numeric using
pandas.to_numeric(..., errors='coerce') and then cast to the nullable Integer
dtype .astype("Int64") so missing/non-numeric values become <NA> instead of
strings; update the assignment for video_data["views"] and
video_data["like_count"] to build Series from the lists, apply to_numeric with
errors='coerce', and then .astype("Int64").

Comment on lines 171 to 173
video_data["text"] = pd.DataFrame(
video_data["title"] + ". " + video_data["description"]
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick (assertive)

Assign Series, not DataFrame, when building text

Avoid wrapping in DataFrame; also coerce to str to be safe.

-            video_data["text"] = pd.DataFrame(
-                video_data["title"] + ". " + video_data["description"]
-            )
+            video_data["text"] = video_data["title"].astype(str) + ". " + video_data["description"].astype(str)
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
video_data["text"] = pd.DataFrame(
video_data["title"] + ". " + video_data["description"]
)
video_data["text"] = video_data["title"].astype(str) + ". " + video_data["description"].astype(str)
🤖 Prompt for AI Agents
In
coursemapper-kg/recommendation/app/services/course_materials/recommendation/youtube_service.py
around lines 171 to 173, the code wraps the concatenated title and description
in a DataFrame; instead assign a pandas Series and ensure both title and
description are coerced to strings. Replace the DataFrame construction with a
Series (e.g., pd.Series((video_data["title"].astype(str) + ". " +
video_data["description"].astype(str)))) so text is a Series of strings rather
than a DataFrame.

Comment on lines +180 to +186
def get_video_details(self, youtube_api_sinlge, video_id):
# print("get_video_details for id -------------------- ", video_id)
try:
duration = (
r["items"][0]["contentDetails"]["duration"]
if r["items"][0]["contentDetails"]["duration"]
else 0
)
views = (
r["items"][0]["statistics"]["viewCount"]
if r["items"][0]["statistics"]["viewCount"]
else 0
)
description = (
r["items"][0]["snippet"]["description"]
if r["items"][0]["snippet"]["description"]
else ""
)

except Exception as e:
print("---------------------------------------")
print(e)
# The number of views are not present for some videos and this leads to an exception. For this
# reason a default value of 0 views will be given that video.
views = 0
duration = (
r["items"][0]["contentDetails"]["duration"]
if r["items"][0]["contentDetails"]["duration"]
else 0
)
description = (
r["items"][0]["snippet"]["description"]
if r["items"][0]["snippet"]["description"]
else ""
r = (
# self.youtube.videos()
youtube_api_sinlge.videos()
.list(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick (assertive)

Typo and minimize API fields

Fix param name and narrow fields to what you actually read.

-    def get_video_details(self, youtube_api_sinlge, video_id):
+    def get_video_details(self, youtube_api_single, video_id):
@@
-                youtube_api_sinlge.videos()
+                youtube_api_single.videos()
                 .list(
-                    part="snippet,statistics,contentDetails",
+                    part="snippet,statistics,contentDetails",
                     id=video_id,
-                    fields="items(statistics,contentDetails(duration),snippet)",
+                    fields="items(statistics(viewCount,likeCount),contentDetails(duration),snippet(description,channelTitle))",
                 )
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
def get_video_details(self, youtube_api_sinlge, video_id):
# print("get_video_details for id -------------------- ", video_id)
try:
duration = (
r["items"][0]["contentDetails"]["duration"]
if r["items"][0]["contentDetails"]["duration"]
else 0
)
views = (
r["items"][0]["statistics"]["viewCount"]
if r["items"][0]["statistics"]["viewCount"]
else 0
)
description = (
r["items"][0]["snippet"]["description"]
if r["items"][0]["snippet"]["description"]
else ""
)
except Exception as e:
print("---------------------------------------")
print(e)
# The number of views are not present for some videos and this leads to an exception. For this
# reason a default value of 0 views will be given that video.
views = 0
duration = (
r["items"][0]["contentDetails"]["duration"]
if r["items"][0]["contentDetails"]["duration"]
else 0
)
description = (
r["items"][0]["snippet"]["description"]
if r["items"][0]["snippet"]["description"]
else ""
r = (
# self.youtube.videos()
youtube_api_sinlge.videos()
.list(
def get_video_details(self, youtube_api_single, video_id):
# print("get_video_details for id -------------------- ", video_id)
try:
r = (
# self.youtube.videos()
youtube_api_single.videos()
.list(
part="snippet,statistics,contentDetails",
id=video_id,
fields="items(statistics(viewCount,likeCount),contentDetails(duration),snippet(description,channelTitle))",
)
)
# …rest of method…

Comment on lines +210 to +217
# Safe gets with defaults (keep same return semantics)
duration = cd.get("duration") if cd.get("duration") else 0
views = st.get("viewCount") if st.get("viewCount") else 0
description = sn.get("description") if sn.get("description") else ""
like_count = st.get("likeCount") if st.get("likeCount") else 0
channel_title = sn.get("channelTitle") if sn.get("channelTitle") else ""

return duration, views, description, like_count, channel_title

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🛠️ Refactor suggestion

Return typed values; default safely

Coerce counts to int and keep empty strings for text fields.

-        duration = cd.get("duration") if cd.get("duration") else 0
-        views = st.get("viewCount") if st.get("viewCount") else 0
-        description = sn.get("description") if sn.get("description") else ""
-        like_count = st.get("likeCount") if st.get("likeCount") else 0
-        channel_title = sn.get("channelTitle") if sn.get("channelTitle") else ""
+        duration = cd.get("duration") or 0
+        views = int(st.get("viewCount") or 0)
+        description = sn.get("description") or ""
+        like_count = int(st.get("likeCount") or 0)
+        channel_title = sn.get("channelTitle") or ""
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
# Safe gets with defaults (keep same return semantics)
duration = cd.get("duration") if cd.get("duration") else 0
views = st.get("viewCount") if st.get("viewCount") else 0
description = sn.get("description") if sn.get("description") else ""
like_count = st.get("likeCount") if st.get("likeCount") else 0
channel_title = sn.get("channelTitle") if sn.get("channelTitle") else ""
return duration, views, description, like_count, channel_title
# Safe gets with defaults (keep same return semantics)
duration = cd.get("duration") or 0
views = int(st.get("viewCount") or 0)
description = sn.get("description") or ""
like_count = int(st.get("likeCount") or 0)
channel_title = sn.get("channelTitle") or ""
return duration, views, description, like_count, channel_title
🤖 Prompt for AI Agents
In
coursemapper-kg/recommendation/app/services/course_materials/recommendation/youtube_service.py
around lines 210 to 217, the numeric fields are returned as-is and may be
strings or None; coerce the count/duration fields to ints and preserve empty
strings for text fields by using safe defaults. Replace the current assignments
with expressions that use the .get(... ) or fallback and coerce to int (e.g.
int(cd.get("duration") or 0), int(st.get("viewCount") or 0),
int(st.get("likeCount") or 0)) and keep description and channel_title as
sn.get("description") or "" and sn.get("channelTitle") or "" so return types are
int, int, str, int, str.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (2)
coursemapper-kg/recommendation/app/services/course_materials/GCN/gcn.py (2)

22-24: Don't disable encryption for Neo4j in production.

encrypted=False exposes data in transit. Make it configurable and default to True (or use a neo4j+s URI).

Apply:

-        self.driver = GraphDatabase.driver(neo4j_uri,
-                                           auth=(neo4j_user, neo4j_pass),
-                                           encrypted=False)
+        self.driver = GraphDatabase.driver(
+            neo4j_uri,
+            auth=(neo4j_user, neo4j_pass),
+            encrypted=getattr(Config, "NEO4J_ENCRYPTED", True),
+        )

87-97: Batch Neo4j writes via UNWIND to avoid N round-trips.

One query per node will be slow at scale.

Apply:

-        with self.driver.session() as session:
-            for i in range(final_embeddings.shape[0]):
-                id = idx[i]
-                f_embedding = final_embeddings[i]
-                embedding = ",".join(str(i) for i in f_embedding)
-                # Find a node in neo4j by its original id and save its final embedding into its "final_embedding" property
-                result = session.run("""MATCH (n) WHERE n.cid= $id or n.sid= $id
-                        set n.final_embedding = $embedding RETURN n""",
-                    id=id,
-                    embedding=embedding)
+        rows = [
+            {"id": str(idx[i]), "embedding": ",".join(map(str, final_embeddings[i]))}
+            for i in range(final_embeddings.shape[0])
+        ]
+        with self.driver.session() as session:
+            session.run(
+                """
+                UNWIND $rows AS row
+                MATCH (n) WHERE n.cid = row.id OR n.sid = row.id
+                SET n.final_embedding = row.embedding
+                """,
+                rows=rows,
+            )
♻️ Duplicate comments (1)
coursemapper-kg/recommendation/app/services/course_materials/GCN/gcn.py (1)

37-37: Prefer debug-level, shape-only logs over commenting out.

Same feedback as before; keep lightweight observability.

Apply:

-        # logger.info(idx_features.shape[0])
+        logger.debug("idx_features count: %d", idx_features.shape[0])
@@
-        # logger.info(features.A)
+        logger.debug("features shape: %s", features.shape)
@@
-        # logger.info(adj.A)
+        logger.debug("adj shape: %s", adj.shape)
@@
-        # logger.info(final_embeddings)
+        logger.debug("final_embeddings shape: %s", final_embeddings.shape)

Also applies to: 41-41, 74-74, 83-83

📜 Review details

Configuration used: CodeRabbit UI

Review profile: ASSERTIVE

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 4b97a54 and 508ee8c.

📒 Files selected for processing (1)
  • coursemapper-kg/recommendation/app/services/course_materials/GCN/gcn.py (4 hunks)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (5)
  • GitHub Check: Analyze (javascript-typescript)
  • GitHub Check: Build container images / Build image (coursemapper-kg/preprocess)
  • GitHub Check: Build container images / Build image (webapp)
  • GitHub Check: Build container images / Build image (coursemapper-kg/concept-map)
  • GitHub Check: Build container images / Build image (coursemapper-kg/recommendation)
🔇 Additional comments (2)
coursemapper-kg/recommendation/app/services/course_materials/GCN/gcn.py (2)

69-69: No-op formatting change.

Safe to keep or drop per project style.


126-127: LGTM; where= supported since NumPy 1.17 and pinned version ==1.26.4 satisfies the requirement.

@@ -66,11 +66,12 @@ def load_data(self):
dtype=np.float32,
)
adj = np.around(adj, 2)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue

np.around densifies sparse matrices — major perf/memory bug.

Rounding the whole sparse matrix converts it to dense; use .data instead.

Apply:

-        adj = np.around(adj, 2)
+        # Round only nonzero weights to keep sparsity
+        adj = adj.tocoo()
+        adj.data = np.around(adj.data, 2)
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
adj = np.around(adj, 2)
# Round only nonzero weights to keep sparsity
adj = adj.tocoo()
adj.data = np.around(adj.data, 2)
🤖 Prompt for AI Agents
In coursemapper-kg/recommendation/app/services/course_materials/GCN/gcn.py
around line 68, calling np.around on the entire sparse adjacency matrix
densifies it and causes major perf/memory issues; instead apply rounding only to
the sparse data array (the .data attribute) and keep the matrix as a sparse
type, i.e., update the matrix's internal data with rounded values and do not
convert or recreate a dense matrix.

# matrix plus its unit matrix and transpose matrix to obtain the complete adjacency matrix
adj = adj + adj.T.multiply(adj.T > adj) - adj.multiply(adj.T > adj)

adj = self.normalize(adj) + sp.eye(adj.shape[0])

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🛠️ Refactor suggestion

Add self-loops before normalization (GCN convention).

Current order changes weights vs. Kipf & Welling’s  = D̂^{-1/2}(A+I)D̂^{-1/2}.

Apply:

-        adj = self.normalize(adj) + sp.eye(adj.shape[0])
+        # Add self-loops, then normalize
+        adj = self.normalize(adj + sp.eye(adj.shape[0]))
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
adj = self.normalize(adj) + sp.eye(adj.shape[0])
# Add self-loops, then normalize
adj = self.normalize(adj + sp.eye(adj.shape[0]))
🤖 Prompt for AI Agents
In coursemapper-kg/recommendation/app/services/course_materials/GCN/gcn.py
around line 73 the code adds self-loops after calling normalize which changes
the resulting weights versus the standard GCN convention; move the addition of
the identity so you add self-loops to adj before calling self.normalize (i.e.,
compute adj_with_loops = adj + sp.eye(adj.shape[0]) then call
self.normalize(adj_with_loops)), and ensure the normalize function computes
degrees and performs symmetric normalization on that augmented matrix.

dependabot Bot and others added 4 commits April 14, 2026 07:53
Bumps nginxinc/nginx-unprivileged from 1.29.1-alpine to 1.29.3-alpine.

---
updated-dependencies:
- dependency-name: nginxinc/nginx-unprivileged
  dependency-version: 1.29.3-alpine
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Bumps nginxinc/nginx-unprivileged from 1.29.0-alpine to 1.29.3-alpine.

---
updated-dependencies:
- dependency-name: nginxinc/nginx-unprivileged
  dependency-version: 1.29.3-alpine
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Bumps [actions/checkout](https://github.com/actions/checkout) from 4 to 6.
- [Release notes](https://github.com/actions/checkout/releases)
- [Changelog](https://github.com/actions/checkout/blob/main/CHANGELOG.md)
- [Commits](actions/checkout@v4...v6)

---
updated-dependencies:
- dependency-name: actions/checkout
  dependency-version: '6'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Update user agent in the response of search query + DNS API

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 4

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (3)
webapp/Dockerfile (1)

20-32: ⚠️ Potential issue | 🟡 Minor

Add HEALTHCHECK to the runtime image.

The final NGINX stage exposes port 4200 but has no healthcheck, reducing failure detection in container orchestration.

Suggested patch
 FROM nginxinc/nginx-unprivileged:1.29.3-alpine
@@
 ENV NGINX_ENTRYPOINT_QUIET_LOGS=1
+HEALTHCHECK --interval=30s --timeout=3s --start-period=10s --retries=3 \
+  CMD wget -q -O /dev/null http://127.0.0.1:4200/ || exit 1
 EXPOSE 4200
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@webapp/Dockerfile` around lines 20 - 32, The Dockerfile final stage (FROM
nginxinc/nginx-unprivileged:1.29.3-alpine) exposes port 4200 but lacks a
HEALTHCHECK; add a HEALTHCHECK instruction after ENV NGINX_ENTRYPOINT_QUIET_LOGS
and before or after EXPOSE 4200 that probes the running nginx (for example an
HTTP GET against localhost:4200 or the nginx status endpoint) with sensible
interval/retries/timeout settings so the container runtime can detect unhealthy
instances; ensure the check runs as the unprivileged user (USER 101) or uses a
simple curl/wget shell command available in the image, and reference the
existing ENV NGINX_ENTRYPOINT_QUIET_LOGS and EXPOSE 4200 in the commit so
reviewers can locate the change.
proxy/Dockerfile (1)

1-7: ⚠️ Potential issue | 🟡 Minor

Add a container HEALTHCHECK for runtime reliability.

There is no healthcheck in this image, so orchestrators cannot detect unhealthy NGINX workers.

Suggested patch
 FROM nginxinc/nginx-unprivileged:1.29.3-alpine
 
 USER root
 COPY --link ./nginx/conf.d/* /etc/nginx/conf.d/
 
+HEALTHCHECK --interval=30s --timeout=3s --start-period=10s --retries=3 \
+  CMD wget -q -O /dev/null http://127.0.0.1:8000/ || exit 1
+
 EXPOSE 8000
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@proxy/Dockerfile` around lines 1 - 7, Add a Docker HEALTHCHECK to the
Dockerfile so orchestrators can detect unhealthy NGINX workers: update the
Dockerfile (the image line FROM nginxinc/nginx-unprivileged:1.29.3-alpine / COPY
and EXPOSE 8000 context) to include a HEALTHCHECK instruction (using CMD-SHELL)
that probes localhost:8000 (e.g., curl/wget --fail or ncat) with sensible
parameters (interval, timeout, retries); place it after the COPY/EXPOSE lines
and ensure it exits non-zero on failure so the container is marked unhealthy.
coursemapper-kg/concept-map/src/services/annotation.py (1)

11-29: ⚠️ Potential issue | 🟠 Major

Inconsistent DBpedia Spotlight endpoint configuration across services.

The concept-map service uses environment variable configuration with a DNS override fallback mechanism, but the recommendation service modules hardcode the same DBpedia Spotlight endpoint in four separate locations without any configuration or DNS handling:

  1. coursemapper-kg/recommendation/app/services/course_materials/kwp_extraction/dbpedia/dataAvailability.py
  2. coursemapper-kg/recommendation/app/services/course_materials/kwp_extraction/dbpedia/concept_tagging.py
  3. coursemapper-kg/recommendation/app/services/course_materials/kwp_extraction/dbpedia/concept_tagging_top_down.py
  4. coursemapper-kg/recommendation/app/services/course_materials/kwp_extraction/dbpedia/concept_tagging Paul.py

All use:

self.url = "https://api.dbpedia-spotlight.org/%s/annotate" % lang

Meanwhile, concept-map uses Config.DBPEDIA_SPOTLIGHT_URL (from environment) and includes a DNS override mechanism to resolve api.dbpedia-spotlight.org to 134.155.98.34.

Impact: This creates a configuration management burden—any endpoint change requires editing multiple hardcoded locations in recommendation services, while concept-map uses environment configuration. The DNS override mechanism in concept-map suggests a specific deployment requirement (firewall, load balancer, specific IP) that recommendation services cannot benefit from.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@coursemapper-kg/concept-map/src/services/annotation.py` around lines 11 - 29,
The recommendation modules hardcode the DBpedia Spotlight endpoint in multiple
places (self.url = "https://api.dbpedia-spotlight.org/%s/annotate" % lang)
instead of using the centralized configuration and DNS override used in
concept-map; update each recommendation class that sets self.url (in
dataAvailability.py, concept_tagging.py, concept_tagging_top_down.py, and
concept_tagging Paul.py) to read the endpoint from the shared configuration (use
Config.DBPEDIA_SPOTLIGHT_URL or an equivalent environment-backed config value)
and remove the hardcoded literal, and ensure any deployment-specific DNS
override logic (the dns_cache/override_dns/new_getaddrinfo mechanism) is
provided centrally at application startup so the recommendation services inherit
the same DNS behavior rather than implementing ad-hoc URLs.
♻️ Duplicate comments (1)
proxy/Dockerfile (1)

2-2: ⚠️ Potential issue | 🟠 Major

Pin the NGINX image to a digest, not a floating tag.

Line 2 still uses a mutable tag, which weakens reproducibility and supply-chain control.

#!/bin/bash
# Verify and retrieve the digest for nginxinc/nginx-unprivileged:1.29.3-alpine
set -euo pipefail

token="$(curl -fsSL 'https://auth.docker.io/token?service=registry.docker.io&scope=repository:nginxinc/nginx-unprivileged:pull' | jq -r '.token')"

curl -fsSI \
  -H "Authorization: Bearer ${token}" \
  -H "Accept: application/vnd.docker.distribution.manifest.v2+json" \
  "https://registry-1.docker.io/v2/nginxinc/nginx-unprivileged/manifests/1.29.3-alpine" \
  | awk 'BEGIN{IGNORECASE=1} /docker-content-digest/ {print $0}'

Expected result: a docker-content-digest header you can pin in FROM nginxinc/nginx-unprivileged@sha256:....

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@proxy/Dockerfile` at line 2, Replace the floating image tag in the
Dockerfile's FROM instruction (currently "FROM
nginxinc/nginx-unprivileged:1.29.3-alpine") with the corresponding immutable
digest; obtain the image digest using the Docker registry manifest endpoint (or
the provided verification snippet) and update the line to use "FROM
nginxinc/nginx-unprivileged@sha256:..." so the build is pinned to the specific
content-addressable image.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In @.github/workflows/release.yml:
- Line 14: The workflow uses the mutable tag actions/checkout@v6 which is a
supply-chain risk; replace that tag with the full immutable commit SHA for the
v6 release (i.e., change uses: actions/checkout@v6 to uses:
actions/checkout@<FULL_COMMIT_SHA>) and add a trailing comment with the
human-friendly tag/version (e.g., # v6.x.y) for maintainability so reviewers can
see which release the SHA corresponds to.

In `@coursemapper-kg/concept-map/src/services/annotation.py`:
- Line 29: Apply the DNS override consistently by centralizing the DBpedia
Spotlight endpoint: create a shared configuration (e.g.,
DBPEDIA_SPOTLIGHT_HOST/DBPEDIA_SPOTLIGHT_URL) or shared init module that calls
override_dns('api.dbpedia-spotlight.org','134.155.98.34') and export the
canonical endpoint, then update annotation.py to import that shared init
(instead of calling override_dns inline) and change the other service that
hardcodes "https://api.dbpedia-spotlight.org/%s/annotate" (see
course_materials/.../kwp_extraction/dbpedia/concept_tagging.py) to use the
shared DBPEDIA_SPOTLIGHT_URL or host variable so all services (including
AnnotationService) resolve via the same override/config.

In `@webapp/Dockerfile`:
- Line 2: Replace the mutable base image tags in the Dockerfile with the
provided immutable digests: update the build stage FROM reference (currently
"node:24.0-slim" in the FROM line used by the build stage) to the pinned digest
"node:24.0-slim@sha256:083430e81f23ca4f309c6de17614d20706ddd544b2adc71fb9fdd86e2371360a",
and update the final runtime FROM reference (currently
"nginxinc/nginx-unprivileged:1.29.3-alpine") to the pinned digest
"nginxinc/nginx-unprivileged:1.29.3-alpine@sha256:5aea7cc516b419e3526f47dd1531be31a56a046cfe44754d94f9383e13e2ee99"
so rebuilds are deterministic.

In `@webserver/src/controllers/knowledgeGraph.controller.js`:
- Around line 515-522: Replace the dead commented axios call and update the
User-Agent header used in the axios.get request inside the code that performs
the Wikimedia API fetch: remove the leftover "// const response = await
axios.get(url);" line and change "User-Agent": "CourseMapper
(coursemapper@example.com)" to a real contact string (project URL or maintainer
email) per Wikimedia policy (e.g., "CourseMapper/1.0
(+https://yourproject.example.com)" or include a real maintainer email) in the
axios.get call so Wikimedia can contact the operator.

---

Outside diff comments:
In `@coursemapper-kg/concept-map/src/services/annotation.py`:
- Around line 11-29: The recommendation modules hardcode the DBpedia Spotlight
endpoint in multiple places (self.url =
"https://api.dbpedia-spotlight.org/%s/annotate" % lang) instead of using the
centralized configuration and DNS override used in concept-map; update each
recommendation class that sets self.url (in dataAvailability.py,
concept_tagging.py, concept_tagging_top_down.py, and concept_tagging Paul.py) to
read the endpoint from the shared configuration (use
Config.DBPEDIA_SPOTLIGHT_URL or an equivalent environment-backed config value)
and remove the hardcoded literal, and ensure any deployment-specific DNS
override logic (the dns_cache/override_dns/new_getaddrinfo mechanism) is
provided centrally at application startup so the recommendation services inherit
the same DNS behavior rather than implementing ad-hoc URLs.

In `@proxy/Dockerfile`:
- Around line 1-7: Add a Docker HEALTHCHECK to the Dockerfile so orchestrators
can detect unhealthy NGINX workers: update the Dockerfile (the image line FROM
nginxinc/nginx-unprivileged:1.29.3-alpine / COPY and EXPOSE 8000 context) to
include a HEALTHCHECK instruction (using CMD-SHELL) that probes localhost:8000
(e.g., curl/wget --fail or ncat) with sensible parameters (interval, timeout,
retries); place it after the COPY/EXPOSE lines and ensure it exits non-zero on
failure so the container is marked unhealthy.

In `@webapp/Dockerfile`:
- Around line 20-32: The Dockerfile final stage (FROM
nginxinc/nginx-unprivileged:1.29.3-alpine) exposes port 4200 but lacks a
HEALTHCHECK; add a HEALTHCHECK instruction after ENV NGINX_ENTRYPOINT_QUIET_LOGS
and before or after EXPOSE 4200 that probes the running nginx (for example an
HTTP GET against localhost:4200 or the nginx status endpoint) with sensible
interval/retries/timeout settings so the container runtime can detect unhealthy
instances; ensure the check runs as the unprivileged user (USER 101) or uses a
simple curl/wget shell command available in the image, and reference the
existing ENV NGINX_ENTRYPOINT_QUIET_LOGS and EXPOSE 4200 in the commit so
reviewers can locate the change.

---

Duplicate comments:
In `@proxy/Dockerfile`:
- Line 2: Replace the floating image tag in the Dockerfile's FROM instruction
(currently "FROM nginxinc/nginx-unprivileged:1.29.3-alpine") with the
corresponding immutable digest; obtain the image digest using the Docker
registry manifest endpoint (or the provided verification snippet) and update the
line to use "FROM nginxinc/nginx-unprivileged@sha256:..." so the build is pinned
to the specific content-addressable image.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 2bc197af-7679-4e4f-bf78-3d51846d0e78

📥 Commits

Reviewing files that changed from the base of the PR and between 508ee8c and 2e9eb15.

📒 Files selected for processing (7)
  • .github/workflows/build.yml
  • .github/workflows/release.yml
  • coursemapper-kg/concept-map/src/graph_db.py
  • coursemapper-kg/concept-map/src/services/annotation.py
  • proxy/Dockerfile
  • webapp/Dockerfile
  • webserver/src/controllers/knowledgeGraph.controller.js

Comment thread .github/workflows/release.yml Outdated
steps:
- name: Checkout repository
uses: actions/checkout@v4
uses: actions/checkout@v6

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Verify that v6 exists and retrieve the exact commit SHA to pin.
# Read-only commands only.

set -euo pipefail

echo "Latest checkout tags (top 20):"
curl -fsSL "https://api.github.com/repos/actions/checkout/tags?per_page=20" | jq -r '.[].name'

echo
echo "Resolve v6 to commit SHA:"
curl -fsSL "https://api.github.com/repos/actions/checkout/git/ref/tags/v6" | jq -r '.object.sha'

Repository: ude-soco/CourseMapper-webserver

Length of output: 172


🏁 Script executed:

#!/bin/bash

# Check if the workflow file exists and examine line 14 and surrounding context
if [ -f ".github/workflows/release.yml" ]; then
  echo "=== Workflow file found ==="
  echo "Lines 1-20 of .github/workflows/release.yml:"
  head -20 ".github/workflows/release.yml" | cat -n
  echo
  echo "=== Full file (checking for permissions declarations) ==="
  cat ".github/workflows/release.yml"
else
  echo "Workflow file not found at .github/workflows/release.yml"
  # List what's in .github/workflows if it exists
  if [ -d ".github/workflows" ]; then
    echo "Contents of .github/workflows/:"
    ls -la ".github/workflows/"
  fi
fi

Repository: ude-soco/CourseMapper-webserver

Length of output: 1626


🌐 Web query:

actions/checkout v6 GitHub Actions tag

💡 Result:

The GitHub Actions actions/checkout v6 tag refers to the major version 6 of the official checkout action, used to check out repositories in workflows. The latest release under v6 is v6.0.2, released on January 9, 2026. Key details: - Tag: actions/checkout@v6 or specifically actions/checkout@v6.0.2 - Main changes in v6.0.0: Improved credential security by storing persist-credentials in $RUNNER_TEMP instead of .git/config (requires Actions Runner v2.329.0+ for Docker containers). No workflow changes needed for most users; git commands like fetch/push still work. - Usage example: - uses: actions/checkout@v6 with: repository: 'owner/repo' # optional ref: 'branch-or-tag' # optional - Requires Node.js 24 runtime support. - Official repository: https://github.com/actions/checkout (50 releases total, latest v6.0.2). As of 2026-04-21, v6 is the current major version.

Citations:


🌐 Web query:

GitHub Actions security best practices pin commit SHA mutable version tags

💡 Result:

GitHub Actions security best practices emphasize pinning actions to full-length commit SHAs instead of mutable version tags (like @v4 or @main) to prevent supply chain attacks where tags can be force-pushed to malicious commits. Key practices: - Use full 40-character SHA: - uses: actions/checkout@11bd719 # v4.2.2 Avoid: - uses: actions/checkout@v4 (mutable tag) - Verify SHA from original repository, not forks. - Add version comment for readability. - Prioritize high-risk/third-party actions. - Use Dependabot/Renovate for automated SHA updates on new releases. - Enable organization/enterprise policies to enforce SHA pinning (fails unpinned workflows). - Pin sub-actions in custom actions. - Combine with least-privilege GITHUB_TOKEN, OIDC for cloud access, code scanning. This ensures immutability: SHAs cannot change without content alteration, mitigating risks shown in real attacks like tj-actions/changed-files (CVE-2025-30066). Official GitHub Docs: Pinning to full SHA is the only immutable reference method.

Citations:


Pin actions/checkout to an immutable commit SHA, not a floating major tag.

Line 14 uses @v6, a mutable major version tag. In workflows with contents: write permissions, this creates supply-chain risk—tags can be force-pushed to malicious commits. Pin to the full commit SHA of the v6 release and include the version in a comment for maintainability.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In @.github/workflows/release.yml at line 14, The workflow uses the mutable tag
actions/checkout@v6 which is a supply-chain risk; replace that tag with the full
immutable commit SHA for the v6 release (i.e., change uses: actions/checkout@v6
to uses: actions/checkout@<FULL_COMMIT_SHA>) and add a trailing comment with the
human-friendly tag/version (e.g., # v6.x.y) for maintainability so reviewers can
see which release the SHA corresponds to.

socket.getaddrinfo = new_getaddrinfo

override_dns('api.dbpedia-spotlight.org', '134.155.95.34')
override_dns('api.dbpedia-spotlight.org', '134.155.98.34')

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Description: Verify the new DBpedia Spotlight IP and service availability

echo "=== Checking DNS resolution for api.dbpedia-spotlight.org ==="
nslookup api.dbpedia-spotlight.org || dig api.dbpedia-spotlight.org

echo -e "\n=== Testing connectivity to new IP 134.155.98.34 ==="
curl -I --connect-timeout 10 https://134.155.98.34/ 2>&1 || echo "Direct IP connection failed"

echo -e "\n=== Testing DBpedia Spotlight API on new IP ==="
curl -s --connect-timeout 10 \
  -H "Accept: application/json" \
  -H "Host: api.dbpedia-spotlight.org" \
  "https://134.155.98.34/en/annotate?text=Berlin&confidence=0.5&support=20" 2>&1

echo -e "\n=== Testing DBpedia Spotlight API via hostname ==="
curl -s --connect-timeout 10 \
  -H "Accept: application/json" \
  "https://api.dbpedia-spotlight.org/en/annotate?text=Berlin&confidence=0.5&support=20" 2>&1

Repository: ude-soco/CourseMapper-webserver

Length of output: 1644


🏁 Script executed:

# Search for DBpedia Spotlight references in the codebase
rg -i "dbpedia" --type py -B2 -A2

Repository: ude-soco/CourseMapper-webserver

Length of output: 50387


🏁 Script executed:

# Search for other DNS override mechanisms
rg "override_dns|socket\.getaddrinfo" --type py -B2 -A2

Repository: ude-soco/CourseMapper-webserver

Length of output: 1416


🏁 Script executed:

# Look for other services that might use the DBpedia API
fd -e py -type f | xargs rg "dbpedia-spotlight" -l

Repository: ude-soco/CourseMapper-webserver

Length of output: 784


Apply DNS override consistently across all DBpedia Spotlight services.

The IP address change to 134.155.98.34 is functional and the service is accessible. However, the DNS override mechanism is applied only in this service, while other recommendation services in the codebase (e.g., coursemapper-kg/recommendation/app/services/course_materials/kwp_extraction/dbpedia/concept_tagging.py) hardcode the full URL without the DNS override. This creates an inconsistency where different services may resolve to different IPs:

  • AnnotationService (this file): Uses override_dns() to map api.dbpedia-spotlight.org → 134.155.98.34
  • Recommendation services: Hardcode https://api.dbpedia-spotlight.org/%s/annotate with default DNS resolution

Consider either:

  1. Centralizing the DNS override in a shared initialization module imported by all services
  2. Using environment variables or configuration to define the DBpedia Spotlight endpoint consistently across all services
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@coursemapper-kg/concept-map/src/services/annotation.py` at line 29, Apply the
DNS override consistently by centralizing the DBpedia Spotlight endpoint: create
a shared configuration (e.g., DBPEDIA_SPOTLIGHT_HOST/DBPEDIA_SPOTLIGHT_URL) or
shared init module that calls
override_dns('api.dbpedia-spotlight.org','134.155.98.34') and export the
canonical endpoint, then update annotation.py to import that shared init
(instead of calling override_dns inline) and change the other service that
hardcodes "https://api.dbpedia-spotlight.org/%s/annotate" (see
course_materials/.../kwp_extraction/dbpedia/concept_tagging.py) to use the
shared DBPEDIA_SPOTLIGHT_URL or host variable so all services (including
AnnotationService) resolve via the same override/config.

Comment thread webapp/Dockerfile
# syntax=docker/dockerfile:1.5
FROM node:22.1-slim as build
# syntax=docker/dockerfile:1.15
FROM node:24.0-slim as build

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Retrieve digests for both Docker Hub images/tags
set -euo pipefail

fetch_digest () {
  local repo="$1"
  local tag="$2"
  local token
  token="$(curl -fsSL "https://auth.docker.io/token?service=registry.docker.io&scope=repository:${repo}:pull" | jq -r '.token')"
  echo "Digest for ${repo}:${tag}"
  curl -fsSI \
    -H "Authorization: Bearer ${token}" \
    -H "Accept: application/vnd.docker.distribution.manifest.v2+json" \
    "https://registry-1.docker.io/v2/${repo}/manifests/${tag}" \
    | awk 'BEGIN{IGNORECASE=1} /docker-content-digest/ {print $0}'
  echo
}

fetch_digest "library/node" "24.0-slim"
fetch_digest "nginxinc/nginx-unprivileged" "1.29.3-alpine"

Repository: ude-soco/CourseMapper-webserver

Length of output: 355


🏁 Script executed:

cat -n webapp/Dockerfile | head -25

Repository: ude-soco/CourseMapper-webserver

Length of output: 859


Pin both base images to immutable digests.

Lines 2 and 20 use mutable tags (node:24.0-slim, nginxinc/nginx-unprivileged:1.29.3-alpine), which makes rebuilds non-deterministic. Update to:

  • Line 2: FROM node:24.0-slim@sha256:083430e81f23ca4f309c6de17614d20706ddd544b2adc71fb9fdd86e2371360a as build
  • Line 20: FROM nginxinc/nginx-unprivileged:1.29.3-alpine@sha256:5aea7cc516b419e3526f47dd1531be31a56a046cfe44754d94f9383e13e2ee99
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@webapp/Dockerfile` at line 2, Replace the mutable base image tags in the
Dockerfile with the provided immutable digests: update the build stage FROM
reference (currently "node:24.0-slim" in the FROM line used by the build stage)
to the pinned digest
"node:24.0-slim@sha256:083430e81f23ca4f309c6de17614d20706ddd544b2adc71fb9fdd86e2371360a",
and update the final runtime FROM reference (currently
"nginxinc/nginx-unprivileged:1.29.3-alpine") to the pinned digest
"nginxinc/nginx-unprivileged:1.29.3-alpine@sha256:5aea7cc516b419e3526f47dd1531be31a56a046cfe44754d94f9383e13e2ee99"
so rebuilds are deterministic.

Comment on lines +515 to +522
// const response = await axios.get(url);
const response = await axios.get(url, {
headers: {
// Use your app name + version + contact (email or URL)
"User-Agent": "CourseMapper (coursemapper@example.com)"
},
timeout: 10000
});

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Replace placeholder contact in User-Agent and drop the dead comment.

Wikimedia's User-Agent policy requires a legitimate way to contact the operator; coursemapper@example.com is a reserved placeholder domain (RFC 2606) and is non-routable. Requests may be throttled or blocked once traffic grows, and there is no way for Wikimedia ops to reach you. Use a real maintainer/project email or the project URL, and remove the leftover // const response = await axios.get(url); line.

🛠️ Proposed fix
-    // const response = await axios.get(url);
-     const response = await axios.get(url, {
-      headers: {
-        // Use your app name + version + contact (email or URL)
-        "User-Agent": "CourseMapper (coursemapper@example.com)"
-      },
-      timeout: 10000
-    });
+    const response = await axios.get(url, {
+      headers: {
+        // App name + version + contact (email or project URL), per Wikimedia UA policy
+        "User-Agent": "CourseMapper/1.0 (+https://github.com/ude-soco/CourseMapper-webserver)",
+      },
+      timeout: 10000,
+    });
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
// const response = await axios.get(url);
const response = await axios.get(url, {
headers: {
// Use your app name + version + contact (email or URL)
"User-Agent": "CourseMapper (coursemapper@example.com)"
},
timeout: 10000
});
const response = await axios.get(url, {
headers: {
// App name + version + contact (email or project URL), per Wikimedia UA policy
"User-Agent": "CourseMapper/1.0 (+https://github.com/ude-soco/CourseMapper-webserver)",
},
timeout: 10000,
});
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@webserver/src/controllers/knowledgeGraph.controller.js` around lines 515 -
522, Replace the dead commented axios call and update the User-Agent header used
in the axios.get request inside the code that performs the Wikimedia API fetch:
remove the leftover "// const response = await axios.get(url);" line and change
"User-Agent": "CourseMapper (coursemapper@example.com)" to a real contact string
(project URL or maintainer email) per Wikimedia policy (e.g., "CourseMapper/1.0
(+https://yourproject.example.com)" or include a real maintainer email) in the
axios.get call so Wikimedia can contact the operator.

rawaa123 and others added 8 commits April 21, 2026 16:56
Bumps nginxinc/nginx-unprivileged from 1.29.3-alpine to 1.31.0-alpine.

---
updated-dependencies:
- dependency-name: nginxinc/nginx-unprivileged
  dependency-version: 1.31.0-alpine
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Bumps nginxinc/nginx-unprivileged from 1.29.3-alpine to 1.31.0-alpine.

---
updated-dependencies:
- dependency-name: nginxinc/nginx-unprivileged
  dependency-version: 1.31.0-alpine
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Bumps [docker/setup-qemu-action](https://github.com/docker/setup-qemu-action) from 3 to 4.
- [Release notes](https://github.com/docker/setup-qemu-action/releases)
- [Commits](docker/setup-qemu-action@v3...v4)

---
updated-dependencies:
- dependency-name: docker/setup-qemu-action
  dependency-version: '4'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Bumps [docker/login-action](https://github.com/docker/login-action) from 3 to 4.
- [Release notes](https://github.com/docker/login-action/releases)
- [Commits](docker/login-action@v3...v4)

---
updated-dependencies:
- dependency-name: docker/login-action
  dependency-version: '4'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Bumps [docker/setup-buildx-action](https://github.com/docker/setup-buildx-action) from 3 to 4.
- [Release notes](https://github.com/docker/setup-buildx-action/releases)
- [Commits](docker/setup-buildx-action@v3...v4)

---
updated-dependencies:
- dependency-name: docker/setup-buildx-action
  dependency-version: '4'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Bumps [docker/build-push-action](https://github.com/docker/build-push-action) from 6 to 7.
- [Release notes](https://github.com/docker/build-push-action/releases)
- [Commits](docker/build-push-action@v6...v7)

---
updated-dependencies:
- dependency-name: docker/build-push-action
  dependency-version: '7'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Bumps [docker/metadata-action](https://github.com/docker/metadata-action) from 5 to 6.
- [Release notes](https://github.com/docker/metadata-action/releases)
- [Commits](docker/metadata-action@v5...v6)

---
updated-dependencies:
- dependency-name: docker/metadata-action
  dependency-version: '6'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Comment thread .github/workflows/build.yml Dismissed
Comment thread .github/workflows/build.yml Dismissed
Comment thread .github/workflows/build.yml Dismissed
Comment thread .github/workflows/build.yml Dismissed
Comment thread .github/workflows/build.yml Dismissed
dependabot Bot and others added 11 commits June 22, 2026 19:52
Bumps nginxinc/nginx-unprivileged from 1.31.0-alpine to 1.31.2-alpine.

---
updated-dependencies:
- dependency-name: nginxinc/nginx-unprivileged
  dependency-version: 1.31.2-alpine
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Bumps nginxinc/nginx-unprivileged from 1.31.0-alpine to 1.31.2-alpine.

---
updated-dependencies:
- dependency-name: nginxinc/nginx-unprivileged
  dependency-version: 1.31.2-alpine
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Bumps [actions/checkout](https://github.com/actions/checkout) from 6 to 7.
- [Release notes](https://github.com/actions/checkout/releases)
- [Changelog](https://github.com/actions/checkout/blob/main/CHANGELOG.md)
- [Commits](actions/checkout@v6...v7)

---
updated-dependencies:
- dependency-name: actions/checkout
  dependency-version: '7'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Bumps nginxinc/nginx-unprivileged from 1.31.2-alpine to 1.31.3-alpine.

---
updated-dependencies:
- dependency-name: nginxinc/nginx-unprivileged
  dependency-version: 1.31.3-alpine
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Bumps nginxinc/nginx-unprivileged from 1.31.2-alpine to 1.31.3-alpine.

---
updated-dependencies:
- dependency-name: nginxinc/nginx-unprivileged
  dependency-version: 1.31.3-alpine
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
1. hiding interaction with process and output from the UI
2. hiding articles tab from UI and modifying articles parallel processing logic from the backend
3. showing only 5 videos as needed for the user study and removing pagination
the keys are now setup in env and used from there via config.py
Bumps nginxinc/nginx-unprivileged from 1.31.3-alpine to 1.31.5-alpine.

---
updated-dependencies:
- dependency-name: nginxinc/nginx-unprivileged
  dependency-version: 1.31.5-alpine
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Bumps nginxinc/nginx-unprivileged from 1.31.3-alpine to 1.31.5-alpine.

---
updated-dependencies:
- dependency-name: nginxinc/nginx-unprivileged
  dependency-version: 1.31.5-alpine
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
@qurat08
qurat08 merged commit cc829fe into main Sep 30, 2026
21 checks passed
@qurat08
qurat08 deleted the dev branch September 30, 2026 13:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants