Local secondary indexes
Some applications only need to query data using the base table's primary key. However,
there might be situations where an alternative sort key would be helpful. To give your
application a choice of sort keys, you can create one or more local secondary indexes on an
Amazon DynamoDB table and issue Query or Scan requests against these
indexes.
Scenario: Using a Local Secondary Index
As an example, consider the Thread table. This table is useful for an
application such as the AWS discussion forums
DynamoDB stores all of the items with the same partition key value continuously. In this
example, given a particular ForumName, a Query operation could
immediately locate all of the threads for that forum. Within a group of items with the
same partition key value, the items are sorted by sort key value. If the sort key
(Subject) is also provided in the query, DynamoDB can narrow down the
results that are returned—for example, returning all of the threads in the "S3"
forum that have a Subject beginning with the letter "a".
Some requests might require more complex data access patterns. For example:
-
Which forum threads get the most views and replies?
-
Which thread in a particular forum has the largest number of messages?
-
How many threads were posted in a particular forum within a particular time period?
To answer these questions, the Query action would not be sufficient.
Instead, you would have to Scan the entire table. For a table with millions
of items, this would consume a large amount of provisioned read throughput and take a
long time to complete.
However, you can specify one or more local secondary indexes on non-key attributes,
such as Replies or LastPostDateTime.
A local secondary index maintains an alternate sort key for a given partition key value. A local secondary index also contains a copy of some or all of the attributes from its base table. You specify which attributes are projected into the local secondary index when you create the table. The data in a local secondary index is organized by the same partition key as the base table, but with a different sort key. This lets you access data items efficiently across this different dimension. For greater query or scan flexibility, you can create up to five local secondary indexes per table.
Suppose that an application needs to find all of the threads that have been posted
within the last three months in a particular forum. Without a local secondary index, the application
would have to Scan the entire Thread table and discard any
posts that were not within the specified time frame. With a local secondary index, a Query
operation could use LastPostDateTime as a sort key and find the data
quickly.
The following diagram shows a local secondary index named LastPostIndex. Note that the
partition key is the same as that of the Thread table, but the sort key is
LastPostDateTime.
Every local secondary index must meet the following conditions:
-
The partition key is the same as that of its base table.
-
The sort key consists of exactly one scalar attribute.
-
The sort key of the base table is projected into the index, where it acts as a non-key attribute.
In this example, the partition key is ForumName and the sort key of the
local secondary index is LastPostDateTime. In addition, the sort key value from the base
table (in this example, Subject) is projected into the index, but it is not
a part of the index key. If an application needs a list that is based on
ForumName and LastPostDateTime, it can issue a
Query request against LastPostIndex. The query results are
sorted by LastPostDateTime, and can be returned in ascending or descending
order. The query can also apply key conditions, such as returning only items that have a
LastPostDateTime within a particular time span.
Every local secondary index automatically contains the partition and sort keys from its base table; you can optionally project non-key attributes into the index. When you query the index, DynamoDB can retrieve these projected attributes efficiently. When you query a local secondary index, the query can also retrieve attributes that are not projected into the index. DynamoDB automatically fetches these attributes from the base table, but at a greater latency and with higher provisioned throughput costs.
For any local secondary index, you can store up to 10 GB of data per distinct partition key value. This figure includes all of the items in the base table, plus all of the items in the indexes, that have the same partition key value. For more information, see Item collections in Local Secondary Indexes.
Attribute projections
With LastPostIndex, an application could use ForumName and
LastPostDateTime as query criteria. However, to retrieve any additional
attributes, DynamoDB must perform additional read operations against the
Thread table. These extra reads are known as
fetches, and they can increase the total amount of provisioned
throughput required for a query.
Suppose that you wanted to populate a webpage with a list of all the threads in "S3" and the number of replies for each thread, sorted by the last reply date/time beginning with the most recent reply. To populate this list, you would need the following attributes:
-
Subject -
Replies -
LastPostDateTime
The most efficient way to query this data and to avoid fetch operations would be to
project the Replies attribute from the table into the local secondary index, as shown in
this diagram.
A projection is the set of attributes that is copied from a table into a secondary index. The partition key and sort key of the table are always projected into the index; you can project other attributes to support your application's query requirements. When you query an index, Amazon DynamoDB can access any attribute in the projection as if those attributes were in a table of their own.
When you create a secondary index, you need to specify the attributes that will be projected into the index. DynamoDB provides three different options for this:
-
KEYS_ONLY – Each item in the index consists only of the table partition key and sort key values, plus the index key values. The
KEYS_ONLYoption results in the smallest possible secondary index. -
INCLUDE – In addition to the attributes described in
KEYS_ONLY, the secondary index will include other non-key attributes that you specify. -
ALL – The secondary index includes all of the attributes from the source table. Because all of the table data is duplicated in the index, an
ALLprojection results in the largest possible secondary index.
In the previous diagram, the non-key attribute Replies is projected into
LastPostIndex. An application can query LastPostIndex
instead of the full Thread table to populate a webpage with
Subject, Replies, and LastPostDateTime. If
any other non-key attributes are requested, DynamoDB would need to fetch those attributes
from the Thread table.
From an application's point of view, fetching additional attributes from the base table is automatic and transparent, so there is no need to rewrite any application logic. However, such fetching can greatly reduce the performance advantage of using a local secondary index.
When you choose the attributes to project into a local secondary index, you must consider the tradeoff between provisioned throughput costs and storage costs:
-
If you need to access just a few attributes with the lowest possible latency, consider projecting only those attributes into a local secondary index. The smaller the index, the less that it costs to store it, and the less your write costs are. If there are attributes that you occasionally need to fetch, the cost for provisioned throughput might well outweigh the longer-term cost of storing those attributes.
-
If your application frequently accesses some non-key attributes, you should consider projecting those attributes into a local secondary index. The additional storage costs for the local secondary index offset the cost of performing frequent table scans.
-
If you need to access most of the non-key attributes on a frequent basis, you can project these attributes—or even the entire base table— into a local secondary index. This gives you maximum flexibility and lowest provisioned throughput consumption, because no fetching would be required. However, your storage cost would increase, or even double if you are projecting all attributes.
-
If your application needs to query a table infrequently, but must perform many writes or updates against the data in the table, consider projecting KEYS_ONLY. The local secondary index would be of minimal size, but would still be available when needed for query activity.
Creating a Local Secondary Index
To create one or more local secondary indexes on a table, use the LocalSecondaryIndexes
parameter of the CreateTable operation. Local secondary indexes on a table
are created when the table is created. When you delete a table, any local secondary
indexes on that table are also deleted.
You must specify one non-key attribute to act as the sort key of the local secondary index. The
attribute that you choose must be a scalar String, Number, or
Binary. Other scalar types, document types, and set types are not
allowed. For a complete list of data types, see Data types.
Important
For tables with local secondary indexes, there is a 10 GB size limit per partition key value. A table with local secondary indexes can store any number of items, as long as the total size for any one partition key value does not exceed 10 GB. For more information, see Item collection size limit.
You can project attributes of any data type into a local secondary index. This includes scalars, documents, and sets. For a complete list of data types, see Data types.
Reading data from a Local Secondary Index
You can retrieve items from a local secondary index using the Query and Scan
operations. The GetItem and BatchGetItem operations can't be
used on a local secondary index.
Querying a Local Secondary Index
In a DynamoDB table, the combined partition key value and sort key value for each
item must be unique. However, in a local secondary index, the sort key value does not need to be
unique for a given partition key value. If there are multiple items in the local secondary index
that have the same sort key value, a Query operation returns all of the
items that have the same partition key value. In the response, the matching items
are not returned in any particular order.
You can query a local secondary index using either eventually consistent or strongly consistent
reads. To specify which type of consistency you want, use the
ConsistentRead parameter of the Query operation. A
strongly consistent read from a local secondary index always returns the latest updated values. If
the query needs to fetch additional attributes from the base table, those attributes
will be consistent with respect to the index.
Example
Consider the following data returned from a Query that requests
data from the discussion threads in a particular forum.
{ "TableName": "Thread", "IndexName": "LastPostIndex", "ConsistentRead": false, "ProjectionExpression": "Subject, LastPostDateTime, Replies, Tags", "KeyConditionExpression": "ForumName = :v_forum and LastPostDateTime between :v_start and :v_end", "ExpressionAttributeValues": { ":v_start": {"S": "2015-08-31T00:00:00.000Z"}, ":v_end": {"S": "2015-11-31T00:00:00.000Z"}, ":v_forum": {"S": "EC2"} } }
In this query:
-
DynamoDB accesses
LastPostIndex, using theForumNamepartition key to locate the index items for "EC2". All of the index items with this key are stored adjacent to each other for rapid retrieval. -
Within this forum, DynamoDB uses the index to look up the keys that match the specified
LastPostDateTimecondition.