Buscar

Mostrando entradas con la etiqueta mongoDB_4_dev_HomeWork. Mostrar todas las entradas
Mostrando entradas con la etiqueta mongoDB_4_dev_HomeWork. Mostrar todas las entradas

jueves, 4 de diciembre de 2014

MongoDB course for developers. unit 6/8. Application Engineering. Homeworks

Homework 6.1





Which of the following statements are true about MongoDB replication. Check all that apply.

Homework 6.2

Let's suppose you have a five member replica set and want to assure that writes are committed to the journal and are acknowledged by at least 3 nodes before you proceed forward. What would be the appropriate settings for w and j?

Homework 6.3

Which of the following statements are true about choosing and using a shard key:

Homework 6.4

You have a sharded system with three shards and have sharded the collections "grades" in the "test" database across those shards. The output of sh.status() when connected to mongos looks like this:
mongos> sh.status()
--- Sharding Status --- 
  sharding version: { "_id" : 1, "version" : 3 }
  shards:
 {  "_id" : "s0",  "host" : "s0/localhost:37017,localhost:37018,localhost:37019" }
 {  "_id" : "s1",  "host" : "s1/localhost:47017,localhost:47018,localhost:47019" }
 {  "_id" : "s2",  "host" : "s2/localhost:57017,localhost:57018,localhost:57019" }
  databases:
 {  "_id" : "admin",  "partitioned" : false,  "primary" : "config" }
 {  "_id" : "test",  "partitioned" : true,  "primary" : "s0" }
  test.grades chunks:
    s1 4
    s0 4
    s2 4
   { "student_id" : { $minKey : 1 } } -->> { "student_id" : 0 } on : s1 Timestamp(12000, 0) 
   { "student_id" : 0 } -->> { "student_id" : 2640 } on : s0 Timestamp(11000, 1) 
   { "student_id" : 2640 } -->> { "student_id" : 91918 } on : s1 Timestamp(10000, 1) 
   { "student_id" : 91918 } -->> { "student_id" : 176201 } on : s0 Timestamp(4000, 2) 
   { "student_id" : 176201 } -->> { "student_id" : 256639 } on : s2 Timestamp(12000, 1) 
   { "student_id" : 256639 } -->> { "student_id" : 344351 } on : s2 Timestamp(6000, 2) 
   { "student_id" : 344351 } -->> { "student_id" : 424983 } on : s0 Timestamp(7000, 2) 
   { "student_id" : 424983 } -->> { "student_id" : 509266 } on : s1 Timestamp(8000, 2) 
   { "student_id" : 509266 } -->> { "student_id" : 596849 } on : s1 Timestamp(9000, 2) 
   { "student_id" : 596849 } -->> { "student_id" : 772260 } on : s0 Timestamp(10000, 2) 
   { "student_id" : 772260 } -->> { "student_id" : 945802 } on : s2 Timestamp(11000, 2) 
   { "student_id" : 945802 } -->> { "student_id" : { $maxKey : 1 } } on : s2 Timestamp(11000, 3) 
If you ran the query
use test
db.grades.find({'student_id':530289})
Which shards would be involved in answering the query?

Homework 6.5

Create three directories for the three mongod processes. On unix, this could be done as follows:
mkdir -p /data/rs1 /data/rs2 /data/rs3
Now start three mongo instances as follows. Note that are three commands. The browser is probably wrapping them visually.
./mongod --replSet m101 --logpath "1.log" --dbpath /data/rs1 --port 27017 --smallfiles --fork

./mongod --replSet m101 --logpath "2.log" --dbpath /data/rs2 --port 27018 --smallfiles --fork

./mongod --replSet m101 --logpath "3.log" --dbpath /data/rs3 --port 27019 --smallfiles --fork
Now connect to a mongo shell and make sure it comes up
./mongo --port 27017
Now you will create the replica set. Type the following commands into the mongo shell:
config = { _id: "m101", members:[
          { _id : 0, host : "localhost:27017"},
          { _id : 1, host : "localhost:27018"},
          { _id : 2, host : "localhost:27019"} ]
};
rs.initiate(config);
At this point, the replica set should be coming up. You can type
rs.status()
to see the state of replication. 

m101:PRIMARY> rs.status()
{
"set" : "m101",
"date" : ISODate("2014-12-04T18:50:17Z"),
"myState" : 1,
"members" : [
{
"_id" : 0,
"name" : "localhost:27017",
"health" : 1,
"state" : 1,
"stateStr" : "PRIMARY",
"uptime" : 234,
"optime" : Timestamp(1417718826, 1),
"optimeDate" : ISODate("2014-12-04T18:47:06Z"),
"electionTime" : Timestamp(1417718835, 1),
"electionDate" : ISODate("2014-12-04T18:47:15Z"),
"self" : true
},
{
"_id" : 1,
"name" : "localhost:27018",
"health" : 1,
"state" : 2,
"stateStr" : "SECONDARY",
"uptime" : 190,
"optime" : Timestamp(1417718826, 1),
"optimeDate" : ISODate("2014-12-04T18:47:06Z"),
"lastHeartbeat" : ISODate("2014-12-04T18:50:17Z"),
"lastHeartbeatRecv" : ISODate("2014-12-04T18:50:17Z"),
"pingMs" : 0,
"syncingTo" : "localhost:27017"
},
{
"_id" : 2,
"name" : "localhost:27019",
"health" : 1,
"state" : 2,
"stateStr" : "SECONDARY",
"uptime" : 190,
"optime" : Timestamp(1417718826, 1),
"optimeDate" : ISODate("2014-12-04T18:47:06Z"),
"lastHeartbeat" : ISODate("2014-12-04T18:50:17Z"),
"lastHeartbeatRecv" : ISODate("2014-12-04T18:50:17Z"),
"pingMs" : 0,
"syncingTo" : "localhost:27017"
}
],
"ok" : 1
}

domingo, 23 de noviembre de 2014

MongoDB for developers 4/8. Performance. Homeworks

Homework 4.1

Suppose you have a collection with the following indexes:
 
> db.products.getIndexes()
[
 {
  "v" : 1,
  "key" : {
   "_id" : 1
  },
  "ns" : "store.products",
  "name" : "_id_"
 },
 {
  "v" : 1,
  "key" : {
   "sku" : 1
  },
                "unique" : true,
  "ns" : "store.products",
  "name" : "sku_1"
 },
 {
  "v" : 1,
  "key" : {
   "price" : -1
  },
  "ns" : "store.products",
  "name" : "price_-1"
 },
 {
  "v" : 1,
  "key" : {
   "description" : 1
  },
  "ns" : "store.products",
  "name" : "description_1"
 },
 {
  "v" : 1,
  "key" : {
   "category" : 1,
   "brand" : 1
  },
  "ns" : "store.products",
  "name" : "category_1_brand_1"
 },
 {
  "v" : 1,
  "key" : {
   "reviews.author" : 1
  },
  "ns" : "store.products",
  "name" : "reviews.author_1"
 }

 Which of the following queries can utilize an index. Check all that apply.

Homework 4.2

Suppose you have a collection called tweets whose documents contain information about the created_at time of the tweet and the user's followers_count at the time they issued the tweet. What can you infer from the following explain output?
 
db.tweets.find({"user.followers_count":{$gt:1000}}).sort({"created_at" : 1 }).limit(10).skip(5000).explain()
{
        "cursor" : "BtreeCursor created_at_-1 reverse",
        "isMultiKey" : false,
        "n" : 10,
        "nscannedObjects" : 46462,
        "nscanned" : 46462,
        "nscannedObjectsAllPlans" : 49763,
        "nscannedAllPlans" : 49763,
        "scanAndOrder" : false,
        "indexOnly" : false,
        "nYields" : 0,
        "nChunkSkips" : 0,
        "millis" : 205,
        "indexBounds" : {
                "created_at" : [
                        [
                                {
                                        "$minElement" : 1
                                },
                                {
                                        "$maxElement" : 1
                                }
                        ]
                ]
        },
        "server" : "localhost.localdomain:27017"
}
 


Homework 4.3

Making the Blog fastPlease download hw4-3.zip from the Download Handout link to get started. This assignment requires Mongo 2.2 or above.
In this homework assignment you will be adding some indexes to the post collection to make the blog fast.
We have provided the full code for the blog application and you don't need to make any changes, or even run the blog. But you can, for fun.
We are also providing a patriotic (if you are an American) data set for the blog. There are 1000 entries with lots of comments and tags. You must load this dataset to complete the problem.
 
# from the mongo shell
use blog
db.posts.drop()
# from the a mac or PC terminal window
mongoimport -d blog -c posts < posts.json
or
mongoimport --host localhost --port 27017 --db blog --collection posts --file "posts.json" --drop --stopOnError

The blog has been enhanced so that it can also display the top 10 most recent posts by tag. There are hyperlinks from the post tags to the page that displays the 10 most recent blog entries for that tag. (run the blog and it will be obvious)
Your assignment is to make the following blog pages fast:
  • The blog home page
  • The page that displays blog posts by tag (http://localhost:8082/tag/whatever)
  • The page that displays a blog entry by permalink (http://localhost:8082/post/permalink)
By fast, we mean that indexes should be in place to satisfy these queries such that we only need to scan the number of documents we are going to return. To figure out what queries you need to optimize, you can read the blog.py code and see what it does to display those pages. Isolate those queries and use explain to explore.

****************************
    # returns an array of num_posts posts, reverse ordered
    def get_posts(self, num_posts):

##################################################################
        self.posts.ensure_index([ ("date", pymongo.DESCENDING)])
##################################################################        

        cursor = self.posts.find().sort('date', direction=-1).limit(num_posts)
        l = []

        for post in cursor:
            post['date'] = post['date'].strftime("%A, %B %d %Y at %I:%M%p") # fix up date
            if 'tags' not in post:
                post['tags'] = [] # fill it in if its not there already
            if 'comments' not in post:
                post['comments'] = []

            l.append({'title':post['title'], 'body':post['body'], 'post_date':post['date'],
                      'permalink':post['permalink'],
                      'tags':post['tags'],
                      'author':post['author'],
                      'comments':post['comments']})

        return l

    # returns an array of num_posts posts, reverse ordered, filtered by tag
    def get_posts_by_tag(self, tag, num_posts):

##################################################################
        self.posts.ensure_index([ ("tags", pymongo.ASCENDING),("date", pymongo.DESCENDING)])
##################################################################        

        cursor = self.posts.find({'tags':tag}).sort('date', direction=-1).limit(num_posts)
        l = []

        for post in cursor:
            post['date'] = post['date'].strftime("%A, %B %d %Y at %I:%M%p")     # fix up date
            if 'tags' not in post:
                post['tags'] = []           # fill it in if its not there already
            if 'comments' not in post:
                post['comments'] = []

            l.append({'title': post['title'], 'body': post['body'], 'post_date': post['date'],
                      'permalink': post['permalink'],
                      'tags': post['tags'],
                      'author': post['author'],
                      'comments': post['comments']})

        return l

    # find a post corresponding to a particular permalink
    def get_post_by_permalink(self, permalink):

##################################################################
        self.posts.ensure_index([ ("permalink", pymongo.ASCENDING)])
##################################################################        

        post = self.posts.find_one({'permalink': permalink})

        if post is not None:
            # fix up likes values. set to zero if data is not present
            for comment in post['comments']:
                if 'num_likes' not in comment:
                    comment['num_likes'] = 0

            # fix up date
            post['date'] = post['date'].strftime("%A, %B %d %Y at %I:%M%p")

        return post


Once you have added the indexes to make those pages fast run the following.
 
python validate.py

(note that for folks who are using MongoLabs or MongoHQ there are some command line options to validate.py to make it possible to use those services) Now enter the validation code below.


Homework 4.4

In this problem you will analyze a profile log taken from a mongoDB instance. To start, please download sysprofile.json from Download Handout link and import it with the following command:

mongoimport -d m101 -c profile < sysprofile.json
or
mongoimport --host localhost --port 27017 --db m101 --collection profile --file "sysprofile.json" --drop --stopOnError

Now query the profile data, looking for all queries to the students collection in the database school2, sorted in order of decreasing latency.

db.profile.find({"ns" : /school2.students/}).sort({"millis":-1}).limit(1).pretty()

What is the latency of the longest running operation to the collection, in milliseconds?


MongoDB for developers 4/8. Performance

Indexes


Indexes are the most important factor in mongodb performance. By default the data is locked for sequently access.


  • Use indexing in order to find sorted data is faster than not indexing
  • The order of the keys in the index is important to find data fast
  • Because indexes take space on disk and need to be updated every write, it is important not to index for all keys of the document- It is much more efficiently to have indexes for all the most common queries
  • We can use the keys of index in the order of are defined. index (a,b,c) --> index for a, index for a,b but not index for c or index for c.



The optimization that have the greatest impact on the performance of a database is adding appropiate indexes on large collections so that only a small percentatge of queries need to scan the collection

Creating Indexes

  • db.collection.ensureIndex({ camp: order (1 or -1)})
  • db.system.indexes.find()     --> show all the indexes
  • db.collection.getIndexes()   --> show the indexes of a collection
  • db.collection.dropIndex({}) --> drop an index of a collection
Please provide the mongo shell command to add an index to a collection namedstudents, having the index key be class, student_name.

db.students.ensureIndex({class:1, student_name:1})

Multikey Indexes

mongodb supports:  an index on a key with value can be an array
mongodb does not support:  an index with combinations of a key with value of an array and other document's elements.
 
An index with more than one combination of arrays and values than are not arrays:
            ensureIndex({a:1, b:1})
            {a:1,b:1}
            {a:[1,2,3] , b:1}        --> support
            {a:[1,2,3] , b:[4,5,6]} --> not support
 
db.collection.find().explain() --> show ahow the find has been made 
 
Suppose we have a collection foo that has an index created as follows:
db.foo.ensureIndex({a:1, b:1})
Which of the following inserts are valid to this collection?
we can make an index of subparts of an array:
        b: [ {a:1,b:1, c: [1,2,3] }]
Two parallel arrays index are not allowed.

Index Creation option, Unique

Please provide the mongo shell command to add a unique index to the collectionstudents on the keys student_id, class_id.
db.students.ensureIndex({student_id:1, class_id:1}, {unique: true})

Index Creation, Removing Dups

db.collection.ensureIndex( {a:1},{unique: true, dropDups:true}}) --> when the index is created, if it finds a dupplicated document, remove all documents that have this dupplicated key, except one

If you choose the dropDups option when creating a unique index, what will the MongoDB do to documents that conflict with an existing index entry?




Delete them for ever and ever, Amen.

Index Creation, Sparse

To create unique indexes when the indexed key is not present in the document.
    1. {a:1,b:2,c:3}
    2. {a:10,b:5,c:10}
    3. {a:13,b:4}
    4. {a:7,b:23}
In this documents, a Spare index will create a index with the present keys discarding the documents that do not contain the same key. If we want to index for {c:1}, the documents 3 and 4 will not be added to the index
  • db.collection.ensureIndex( {a:1},{unique: true, sparse:true}}) --> it will create a disperse index
  • db.collection.find().sort().hint() --> hint() forces the query optimizer to use a specific index to fulfill the query
Suppose you had the following documents in a collection called people with the following docs:
> db.people.find()
{ "_id" : ObjectId("50a464fb0a9dfcc4f19d6271"), "name" : "Andrew", "title" : "Jester" }
{ "_id" : ObjectId("50a4650c0a9dfcc4f19d6272"), "name" : "Dwight", "title" : "CEO" }
{ "_id" : ObjectId("50a465280a9dfcc4f19d6273"), "name" : "John" }
And there is an index defined as follows:
db.people.ensureIndex({title:1}, {sparse:1})
If you perform the following query, what do you get back, and why?
db.people.find({title:null})
No documents, because the query uses the index and there are no documents with title:null in the index.


Index Creation, Background


foreground (default)    Background:
        faster                     slow
        block writes            dos not block writers
            (per DBlock)


Which things are true about creating an index in the background in MongoDB. Check all that apply.




Using Explain

Inform how the query was done,which index was used to and how they were used.
Given the following output from explain, what is the best description of what happened during the query?
{
 "cursor" : "BasicCursor",
 "isMultiKey" : false,
 "n" : 100000,
 "nscannedObjects" : 10000000,
 "nscanned" : 10000000,
 "nscannedObjectsAllPlans" : 10000000,
 "nscannedAllPlans" : 10000000,
 "scanAndOrder" : false,
 "indexOnly" : false,
 "nYields" : 7,
 "nChunkSkips" : 0,
 "millis" : 5151,
 "indexBounds" : {
  
 },
 "server" : "Andrews-iMac.local:27017"
}
The query scanned 10,000,000 documents, returning 100,000 in 5.2 seconds.


When is an index used?

MongoDb extract estatistic information of the useful queries and choose the best indexation in background every 100 queries more or less.

Given collection foo with the following index:
db.foo.ensureIndex({a:1, b:1, c:1})
Which of the following queries will use the index?

How large is your index?




Indexes have to be in memory in order to get good performance. The size of the index can be very big and will use a lot of memory. This is a consideration at time to planning what sort of indexes we want to create for the documents that we have.
  • db.collection.stats()              --> statistic information
  • db.collection.totalIndexSize() --> get information of size on disc of indexes
Is it more important that your index or your data fit into memory?


Index Cardinality




  • Regular index: 1 to 1
  • Sparse index: <= documents
  • Multikey index: with array of tags  > number of documents
Let's say you update a document with a key called tags and that update causes the document to need to get moved on disk. If the document has 100 tags in it, and if the tags array is indexed with a multikey index, how many index points need to be updated in the index to accomodate the move?
100

Indexing in pyMongo

db.collection.ensureIndex([ ('key1', pymongo.ASCENDING), ('key2', pymongo.DESCENDING)])   

Hinting an Index

  • db.people.find().sort({'title':1}).hint({'title:1}
  • db.people.find().sort({'title':1}).hint({ $natural:1 }) --> specify the index which is the best for mongodb 
hint() specify wich index will be used. Using an index with a key that do not exist in the documents, the query cannot be executed because there is not any pointer in the index to any document.
 
Given the following data in a collection:
> db.people.find()
{ "_id" : ObjectId("50a464fb0a9dfcc4f19d6271"), "name" : "Andrew", "title" : "Jester" }
{ "_id" : ObjectId("50a4650c0a9dfcc4f19d6272"), "name" : "Dwight", "title" : "CEO" }
{ "_id" : ObjectId("50a465280a9dfcc4f19d6273"), "name" : "John" }
and the following indexex:
> db.people.getIndexes()
[
 {
  "v" : 1,
  "key" : {
   "_id" : 1
  },
  "ns" : "test.people",
  "name" : "_id_"
 },
 {
  "v" : 1,
  "key" : {
   "title" : 1
  },
  "ns" : "test.people",
  "name" : "title_1",
  "sparse" : 1
 }
]
Which query below will return the most documents.
hint natural to use BasicCursor returns all docs.

Efficiency of index use

There are elements that $gt, $lt, $eq, $ne, $exist  than can make the query slow because have to examine all the documents.
Is better to use regular expressions /abcd/ -> look for a,b,c,d, /^abcd/ do not look for a,b,c,d
 
Keep in mind when you think aboinut indexing you have to consider how the index was used: only for the sort o if it was used inefficiently and caused that de database examined millions of records, etc

  

Geospatial Indexes




They are indexes based in locations using 2D coordinates. : {'location': [x,y] }
 
ensureIndex({ "location": '2d', type: 1})
find({location: { "$near" : [x,y] }} ) --> retorn locatiosn in increase distances
 
Suppose you have a 2D geospatial index defined on the key location in the collection places. Write a query that will find the closest three places (the closest three documents) to the location 74, 140.
 
db.places.find({location: {$near: [74,140]}}).limit(3)

Geospatial Spherical

  • lng -> vertical 
  • lat  -> horizontal (-90 to 90)
specification GeoJSON -> ( )
     { "location" : { Type : "Point", "coordinates : [-122,40] "} }
 
ensureIndex( { location : '2dsphere'})
 
find( { "location" :
            { "$near" :
                { "$geometry" :
                    { "type"         : "Point" ,
                      "coordinates"  : [-10,10] },
                      "$maxdistante" : 2000 <-- in meters
                     }
                }
            })


What is the query that will query a collection named "stores" to return the stores that are within 1,000,000 meters of the location latitude=39, longitude=-130? Type the query in the box below. Assume the stores collection has a 2dsphere index on "loc" and please use the "$near" operator. Each store record looks like this: 
 
{ "_id" : { "$oid" : "535471aaf28b4d8ee1e1c86f" },
  "store_id" : 8, 
  "loc" : { "type" : "Point", "coordinates" : [ -37.47891236119904, 4.488667018711567 ] } }
 
db.stores.find( { loc : { "$near" : { "$geometry" : { "type" : "Point", "coordinates : [ -130, 39]},"$maxdistance" : 1000000}}})

Full Text searches in mongoDb

There is a type of index that allow to look for text in the data.
 
ensureIndex( { 'words': 'text'})
 
db.collection.find( { "$text" : {"$search":'texto'}) --> look for dog in the documents no case-sensitive.
 
db.collection.find( { "$text" : {"$search":'word1 word2 word3 '}}, { "score" : {"$meta" : 'textScore'}}).sort( { "score": { "$meta" : 'textScore'}}) --> look for documents that contains all of the tree wordsYou create a text index on the "title" field of the movies collection, and then perform the following text search:

> db.movies.find( { $text : { $search : "Big Lebowski" } } )

Which of the following documents will be returned, assuming they are in the movies collection? Check all that apply.

Logging and profiling: log slow queries

Mongodb have a profiler to detect via log informing how mongod is accessing to database:  system.profile
 
We have tree levels to log information of tue queries to know how our application is working:
  • level 0: default and it is log off
  • level 1: only log slow queries (register slow queries)
  • level 2: record all logs of the queries   (register my queries) --> is for debugging
mongod --db dbpath --profile 1 --slows 2 (2 mseconds)
  • db.system.profile.find()
  • db.getProfilingLevel()
  • db.getProfilingStatus()
  • db.setProfilingLevel(1,4)  level 1 , 4 mseconds
  • db.setProfilingLevel(0) --> off
Write the query to look in the system profile collection for all queries that took longer than one second, ordered by timestamp descending.

db.system.profile.find({millis: {$gt: 1000}}).sort({ts: -1})



MongoStat

system information of mongodb database.

    column idx miss% --> % de perdida de memoria por los índices

  

MongoTop

 give a high level view of how mongdb is spending the time.

 

Resume

    1. indexes are critical to performance
    2. explain()
    3. hint()
    4. profiling

 

Sharding

It is a techique to divide a collections in multiples servers
 

Application --> mongos --> mongod 1
                                       --> mongod 2
                                       --> mongod 3
 

It is necessaryt to include a sharding key to look for the server in which document is