Insight Guard
Save
Abstract
The mission of InsightGuard is to harness the power of machine learning and deep learning to develop a cutting-edge suicide prediction tool capable of analyzing text for signs of suicidal ideation. By identifying linguistic patterns and contextual markers in real-time, InsightGuard aims to empower mental health professionals, organizations, and researchers with actionable insights, facilitating timely interventions. The ultimate goal is to bridge the gap between silent suffering and life-saving support, contributing to a more compassionate and proactive approach to mental health care.
Introduction
Mental health issues, particularly suicidal ideation, represent a critical global health challenge. With over 700,000 deaths annually attributed to suicide, the need for timely detection and intervention has never been greater. Traditional methods of identifying individuals at risk often fall short due to stigma, limited accessibility, and delayed responses. InsightGuard addresses this gap by providing a scalable, automated solution that analyzes text data from social media platforms and other sources, delivering real-time insights into potential risks. By leveraging advanced natural language processing (NLP) and machine learning technologies, it offers an innovative approach to identifying suicidal ideation through language patterns. Its user-friendly interface ensures accessibility for a wide range of stakeholders, from mental health professionals to researchers, while maintaining the precision required for effective early intervention.
Background and Motivation
Suicide is a multifaceted issue influenced by psychological, social, and environmental factors. Despite the growing awareness of mental health challenges, many individuals experiencing suicidal thoughts fail to seek help, often expressing their struggles indirectly through online platforms. Social media and other textual communication channels have become critical arenas for understanding and addressing mental health issues. Existing tools for sentiment analysis and mental health monitoring provide valuable insights but lack the specificity and focus required for accurate suicidal ideation detection. This gap motivated the development of InsightGuard, a tool designed to overcome these limitations by tailoring its algorithms to detect subtle linguistic cues associated with mental health distress. The motivation behind InsightGuard stems from a deep commitment to addressing the unmet need for timely and precise tools in suicide prevention. By empowering stakeholders with real-time, actionable insights, the project seeks to transform how mental health issues are identified and addressed, ultimately reducing the burden of suicide on individuals, families, and society.
Dataset
- About the Dataset
- Dataset Characteristics
- Data Quality
About the Dataset
The InsightGuard project aims to address the challenge of detecting suicide ideation through text analysis. During the development of this project, it became evident that there was a lack of publicly available datasets specifically focused on suicide ideation detection. To fill this gap, a comprehensive dataset was created to aid researchers, mental health professionals, and developers in building and enhancing tools for early detection of suicide ideation based on textual data. The dataset is a collection of posts from the Reddit platform, primarily sourced from the SuicideWatch, depression, and r/teenagers subreddits. Data was collected using the Pushshift API, which allows for the retrieval of historical posts from Reddit. The dataset spans from 2008 to 2021, capturing a broad range of discussions related to mental health.
SuicideWatch Posts: Posts from the SuicideWatch subreddit, where users often express suicidal thoughts or seek help. These posts were collected from December 16, 2008, to January 2, 2021, and are labeled as suicide.
Non-Suicidal Posts: To balance the dataset, non-suicidal posts were collected from the r/teenagers subreddit. These posts are unrelated to suicide and depression, offering a non-suicidal comparison group.
This dataset is designed to aid in the development of machine learning models and tools that can detect language patterns indicative of suicidal ideation, ultimately contributing to suicide prevention and mental health awareness initiatives.
Methodology
- Data Collection
- Data Preprocessing
- Feature Extraction
- Model Development
Data Collection
Data is a critical component in training and evaluating machine learning models. For this project, a dataset comprising posts from Reddit's SuicideWatch and depression subreddits was utilized. The data was collected using the Pushshift API, which provided posts from SuicideWatch (from Dec 16, 2008, to Jan 2, 2021) and depression (from Jan 1, 2009, to Jan 2, 2021). Non-suicide posts were collected from the r/teenagers subreddit. This dataset consists of labeled posts, with each post marked as either suicide or non-suicide.
Results
To understand the performance of our models, we employ a set of evaluation metrics that provide insights into their accuracy and reliability.
- Deep Learning
- Logistic Regression
- K-Nearest Neighbors
- Multinomial Naive Bayes
- Random Forest
Deep Learning (A BiLSTM-RNN Architecture)
The model is a deep learning architecture composed of several layers designed to effectively capture the complex patterns and nuances present in textual data related to suicide risk. The architecture comprises the following layers:
-
Embedding Layer
- Maps words or tokens into dense vector representations.
- Utilizes pre-trained word embeddings from Gensim to initialize the weights.
- Ensures that words with similar meanings have similar representations.
-
Multiple Bidirectional LSTM Layers
- Processes input sequences in both forward and backward directions, capturing long-range dependencies.
- Captures context from both past and future, enabling the model to understand the overall sentiment and intent of the text.
-
Simple RNN Layers
- Processes input sequences in a single direction, capturing short-range dependencies.
- Further extracts relevant features from the input sequences.
- Contributes to capturing the overall sentiment and intent of the text.
-
Dense Layer
- Performs a non-linear transformation on the input.
- Maps the extracted features to a higher-level representation.
-
Output Layer
- Employs a softmax activation function to output probabilities for each class (suicide or non-suicide).
Performance
Precision: 0.903
Recall: 0.964
F1 Score: 0.933
Accuracy: 0.930
Meet the Team
Rahul Chhatbar
Computer Science