- Understand the motivation for non-relational data stores
- Be familiar with Cassandra at a high-level
- Be familiar with basic installation / setup of Cassandra, and how an installation is structured
- Understand the Cassandra architecture, including the cluster structure and partitioners
- Understand and use data replication and eventual consistency with Cassandra
- Learn the basics of the Cassandra data model, and how to create good data models
- Use CQL 3 to create tables and execute queries
- Learn and use cqlsh
- Know the standard CQL data types
- Understand and use primary keys, compound primary keys, and composite partition keys
- Understand and use collections, secondary indexes, counters, and batches
- Understand and use Compare and Set (CAS) / Lightweight Transactions
- Be comfortable with and use other Cassandra 2 features such as static columns
- Understand the structure of the Java driver
- Use the basic Java API to connect to and work with Cassandra
- Use QueryBuilder to build dynamic queries
- Use asynchronous queries
Hands-On
Minimum 50% hands-on
Supported Platforms
Cassandra 2.0.6+ and the DataStax Java Driver 2.0.1+ on Linux Operating Systems (VM provided for labs)
Big Data Training Prerequisites
Reasonable Java experience for the Java driver labs, some knowledge of databases
Big Data Training Course Duration
3 Days
Big Data Training Course outline
Session 1: Introduction to Cassandra
- Overview
- The motivation for non-relational data stores
- Why relational databases don't support modern applications well
- Cassandra at a high-level
- Use cases
- Features Strengths (Scalability, robustness, linear performance with scale-out), etc.
- High Level Structure
- Acquiring and Installing Cassandra
- Configuring and Installation Structure
- LABS:
- Configure, Start/Stop Cassandra
- StockWatcher Demo
Session 2: Overview of Architecture and Data Model
- Basic Cassandra Architecture
- Cluster Structure - Nodes, Virtual Nodes, Ring Topology
- Consistent Hashing, Tokens, Partitioners, and Data Distribution
- Data Replication, the Replication Factor, Keyspaces
- Consistency, the CAP theorem, Eventual Consistency
- The C* Data Model
- Data Model and CQL 3 Introduction
- Using CQL and cqlsh
- Single primary key tables and how to define them using CQL
- Inserting Data (INSERT), Data Distribution in the Ring, Upsert
- Querying for Data (SELECT)
- CQL Data Types
- Working with Primary Keys
- LABS:
- Spin up the Lab Cluster
- Create Simple Tables
- Insert/Query Tables
- Use copy to Populate a Table
Session 3: The Cassandra Data Model
- Compound Primary Keys
- CQL table definition
- The partition key and clustering columns
- CQL Mapping vs. Internal Storage View
- Other Capabilities
- Expiring Columns / Time To Live (TTL)
- Batches
- Clustering order, ORDER BY, and CLUSTERING ORDER BY
- Filtering results and ALLOW FILTERING
- Composite Partition Keys
- Motivation and uses
- CQL Definition
- Effect on Partitioning and Internal Storage View
- Indexes and Secondary Indexes
- Partition Key Indexes, token()
- Non-primary Key (Secondary) Indexes
- Guidelines and Querying
- Understand and Use Counters
- Motivation and Uses
- Structure, Characteristics, CQL, Usage
- Limitations
- Understand and use collections
- Motivation and uses
- CQL definition (set, list, and map)
- Inserting, Updating, Deleting with a Collection
- Limitations and Internal Storage View
- LABS:
- Introduce Compound Primary Keys
- Refine Compound Primary Key
- Work with Composite Partition Keys
- Secondary Indexes
- Counters
- Using Collections
Session 4: Data Consistency
- Data Consistency in Cassandra
- Tunable Consistency
- CAP Theorem
- Coordinators and Client Requests
- Consistency Levels in C* - ONE, QUORUM, ALL
- Configuring Immediate Consistency
- CL ONE is Your Friend
- Other CL Levels
- Compare and Set (CAS) / Lightweight Transactions
- Motivation and Need
- Overview of CAS
- Using CAS - IF NOT EXISTS, IF condition
- Paxos - How CAS works
- Overhead and caveats
- Static Columns
- Overview
- Declaring Tables and Using Static Columns
- Guidelines and Uses
- Repair Mechanisms
- Read Repair
- Hinted Handoff
- LABS:
Session 5: How Things Work
- Write Failures
- Unavailability, Node Failure
- Requirements for Writing
- Key and Row Caches
- Cache Overview
- Guidelines
- Multi-Data Center Support
- Overview
- Replication Factor Configuration
- Consistency Levels - LOCAL/EACH QUORUM
- Deletes
- CQL
- Tombstones
- Issues and Guidelines
- LABS:
Session 6: The Java Driver
- Introduction
- Overview and Architecture
- Features
- API Introduction
- Cluster and Cluster.Builder
- Creating the Cluster
- Contact Points
- Getting a Session
- Working with Sessions
- Querying
- PreparedStatement, Statement, BoundStatement
- Using PreparedStatements, Binding Values
- Executing the Query
- Processing Query Results
- CQL to Java Type Mapping
- Working with UUIDs
- Setting Consistency Level
- QueryBuilder and Dynamic Queries
- Dynamic Queries
- Bind Variables and SimpleStatement
- QueryBuilder - Fluent API for Queries
- Building SELECT, Select Type, Select.where(), Chaining WHERE Clauses
- Building DELETE, Delete Type, Delete.where()
- Building INSERT, Insert Type, Insert.insertInto(), Insert.value()
- Building Update, Update Type, Insert.with(), Insert.value()
- Building Regular Batches and Prepared Statement Batches (BatchStatemnt)
- Other Queries
- Asynchronous Querying
- Normal vs Asynchronous Querying
- Interface java.util.concurrent.Future
- ResultSetFuture
- Querying Asynchronously and Processing Future Results
- Listeners for Processing
- Driver Policies
- Overview
- Load Balancing Policies - RoundRobinPolicy, DCAwareRoundRobinPolicy
- Retry Policies - DefaultRetryPolicy, DowngradingConsistencyRetryPolicy, LatencyAwarePolicy
- The Policies Class
- LABS:
- Introducing StockWatcher (the Lab Domain)
- Connect to a Cluster
- Execute Queries
- Use QueryBuilder for Select
- Use QueryBuilder More Extensively
- Use executeAsync() and ResultSetFuture
- Use Driver Policies
- StockWatcher Deep Dive - Examining a Full-fledged Application