Subject

Big Data Modelling and Management

1. Course Title Big Data Modelling and Management
Big data modeling and management
2. Code m23_s_055
3. Study Programme
4. Organizer of the study programme (unit, institute, department or division) Faculty of Computer Science and Engineering
5. Degree level (first, second, third cycle) Second cycle
6. Academic year / semester 10 / Summer
7. Number of ECTS credits 6
8. Teacher Eftim Zdravevski, Goran Velinov
9. Prerequisites for enrolling in the course
10. Objectives of the course programme (competences) The development trends of traditional relational (SQL) database management systems, data warehouses, as well as the concepts of NoSQL and NewSQL big data management systems will be studied. The concepts of storing data on various memory media will be examined. Approaches for centralized or distributed storage, as well as logical organization by rows, columns, graphs, or documents, will be studied. Methods for partitioning and indexing structured, (semi/un)structured, and textual data will be studied.
Real-world approaches and solutions will be covered for overcoming the challenges of modeling, management, implementation, and deployment of big data systems. By the end of the course, students will know which systems are most suitable and what steps are required to introduce big data systems in companies, as well as the challenges companies face.
11. Course content A new perspective on data warehouses: conceptual, logical, and physical models; data lake concepts. Overview of big data management systems. Data modeling in big data systems: implications of time in data modeling.
Concepts of database organizations by columns (MonetDB, HBase, Cassandra), by key-value (DynamoDB, Riak), by documents (MongoDB, CouchDB), and in graphs (Neo4j, OrientDB).
Transactional and analytical databases running in main memory. Alternative data storage media.
Indexing and partitioning strategies and their impact on scalability and performance; Text indexing databases (Solr, Elasticsearch)
Integration of various data sources; Planning for development, capacity, and infrastructure.
Systems and tools for analyzing large static data, such as Spark, Spark SQL, Hive, Pig, Tez, and newer ones.
Techniques for handling data streams; systems and tools for analyzing dynamic big data such as Spark Streaming, Storm, Oozie, Sqoop, Flink, and newer.
Managing System Deployments: Expectations, Assumptions, Risks, and Team Building
Strategies and scenarios for migrations, security, and backup of big data.
12. Learning methods Lectures supported by slide presentations, interactive lectures, practical classes (using equipment and software packages), teamwork, case studies, guest lecturers, independent preparation and defence of a project assignment and seminar paper, and learning in an electronic environment (forums and consultations).
13. Total available time 6 ECTS x 30 hours = 180 hours
14. Distribution of available time 30 + 30 + 30 + 45 + 45 = 180 hours
15. Forms of teaching activities
15.1. Lectures - theoretical instruction 30 hours
15.2. Exercises (laboratory, auditory), seminars, teamwork 30 hours
16. Other forms of activities
16.1. Project assignments 45 hours
16.2. Independent assignments 30 hours
16.3. Home study 45 hours
17. Assessment method
17.1. Tests 30 points
17.2. Seminar paper / project (presentation: written and oral) 45 points
17.3. Activities and learning 20 points
17.4. Final exam 0 points
18. Grading criteria (points / grade)
up to 50 points5 (five) (F)
from 51 to 60 points6 (six) (E)
from 61 to 70 points7 (seven) (D)
from 71 to 80 points8 (eight) (C)
from 81 to 90 points9 (nine) (B)
from 91 to 100 points10 (ten) (A)
19. Requirement for obtaining a signature and taking the final exam completed activities 15.1 and 15.2
20. Language of instruction Macedonian and
21. Method for monitoring the quality of teaching internal evaluation and survey mechanism
22. Literature
22.1. Required literature
1. Franz Faerber, Alfons Kemper, Per-Åke Larson, Justin Levandoski, Thomas Neumann and Andrew Pavlo | Main Memory Database Systems, Foundations and Trends in Databases | Now Publishers | 2017
2. Daniel Abadi, Peter Boncz, Stavros Harizopoulos, Stratos Idreos and Samuel Madden | The Design and Implementation of Modern Column-Oriented Database Systems | Now Publishers | 2013
3. García Márquez, Fausto Pedro, Lev, Benjamin | Big Data Management | Springer | 2017
4. Corea, Francesco | Big Data Analytics: A Management Perspective | Springer | 2016
5. Moshirpour, Mohammad, Far, Behrouz, Alhajj, Reda | Highlighting the Importance of Big Data Management and Analysis for Various Applications | Springer | 2018
6. Sherif Sakr and Mohamed Gaber | Large Scale and Big Data: Processing and Management | CRC Press | 2014
7. Shivnath Babu and Herodotos Herodotou | Massively Parallel Databases and MapReduce Systems | Now Publishers | 2013
22.2. Additional literature
No. Author Title Publisher Year