Project
CS4315/CS6315 Artificial Intelligence and Machine Learning
Machine Learning Projects Template
Develop your own machine learning projects are an excellent way to showcase your skills and knowledge in real-world scenarios. Jupyter Notebooks and Google Colab can be handy tools for trying new ideas and simple experiments. However, for more complex and larger projects, it is recommended to use a structured approach to organize your code, data, and documentation. This will help you maintain clarity and efficiency as your project grows. Below is a popular template that you can use to structure your machine learning projects. This template is adopt from the Cookiecutter Data Science website, which contains some essential parts for organizing your machine learning project effectively.
├── LICENSE <- Open-source license if one is chosen
├── Makefile <- Makefile with convenience commands like "make data" or "make train"
├── README.md <- The top-level README for developers using this project.
├── data
│ ├── external <- Data from third-party sources.
│ ├── interim <- Intermediate data that has been transformed.
│ ├── processed <- The final, canonical data sets for modeling.
│ └── raw <- The original, immutable data dump.
│
├── docs <- A default mkdocs project; see www.mkdocs.org for details
│
├── models <- Trained and serialized models, model predictions, or model summaries
│
├── notebooks <- Jupyter notebooks.
│
├── pyproject.toml <- Project configuration file with package metadata for
│ src and configuration for tools like black
│
├── references <- Data dictionaries, manuals, and all other explanatory materials.
│
├── reports <- Generated analysis as HTML, PDF, LaTeX, etc.
│ └── figures <- Generated graphics and figures to be used in reporting
│
├── requirements.txt <- The requirements file for reproducing the analysis environment, e.g.
│ generated with `pip freeze > requirements.txt`
│
├── setup.cfg <- Configuration file for flake8
│
└── src <- Source code for use in this project.
│
├── __init__.py <- Makes src a Python module
│
├── config.py <- Store useful variables and configuration
│
├── dataset.py <- Scripts to download or generate data
│
├── features.py <- Code to create features for modeling
│
├── modeling
│ ├── __init__.py
│ ├── predict.py <- Code to run model inference with trained models
│ └── train.py <- Code to train models
│
└── plots.py <- Code to create visualizations Expectations for the Final Project
The final project is an opportunity for you to apply the knowledge and skills you have gained throughout the course to a real-world problem. You are expected to choose a dataset, define a problem statement, and implement a machine learning solution. The project should demonstrate your ability to preprocess data, select appropriate models, evaluate model performance, and communicate your findings effectively. You are encouraged to be creative and explore different approaches to solving the problem. The final project will be evaluated based on the following criteria:
- Problem Definition: Clearly define the problem you are trying to solve and explain why it is important
- Data Preprocessing: Demonstrate your ability to clean and preprocess the data effectively
- Model Selection: Justify your choice of machine learning models and explain how they are appropriate for the problem
- Model Evaluation: Evaluate the performance of your models using appropriate metrics and techniques
- Communication: Present your findings in a clear and concise manner, using visualizations and explanations to support your conclusions
For In-Class Students
Please form a group of maximum 3 students (group size can be 1, 2, and 3), and submit a single project report and the code zip file for the group. Each member of the group should contribute to the project and be able to explain their contributions in the report. The final project report should include a detailed description of the problem, data preprocessing steps, model selection and evaluation, and conclusions. The report should be well-organized and clearly written, with good visualizations to support your findings.
Here is an example of the group project member contribution table:
| Student | Primary Responsibilities | Deliverables |
|---|---|---|
| Student A | Data Preprocessing | Clean dataset and preprocessing code |
| Student B | Model Development | Model selection andimplementation |
| (A and B) or C | Model Evaluation | Metrics and figures |
Make sure to clearly define the responsibilities of each group member and ensure that everyone contributes to the project. The final project report should be submitted as a single document, with each member’s contributions clearly indicated. Note that the final project presentation will be OPTIONAL for in-class students on the final scheduled class and it is NOT a Requirement for the final project grading.
For Remote Graduate Students
Graduate students are expected to work on an individual project. The final project report should include a detailed description of the problem, data preprocessing steps, model selection and evaluation, and conclusions. The report should be well-organized and clearly written, with appropriate visualizations to support your findings. The final project report should be submitted as a single document, with a clear explanation of your training methodology and results. Note that the final project presentation will be REQUIRED for remote students on the final scheduled class and it is Mandatory for the final project grading.
Machine Learning Final Project Ideas
The examples provided above can serve as inspiration for your Final Machine Learning Project. You are also welcome to propose your own project ideas, but please note that they will require prior approval from the instructor (discuss with the instructor first).