Buscar

Mostrando entradas con la etiqueta mongoDB_4_dev. Mostrar todas las entradas
Mostrando entradas con la etiqueta mongoDB_4_dev. Mostrar todas las entradas

martes, 13 de enero de 2015

MongoDB course for developers. unit 5/8. Aggregation framework. Homeworks

Homework 5.1


Finding the most frequent author of comments on your blog
In this assignment you will use the aggregation framework to find the most frequent author of comments on your blog. We will be using a data set similar to ones we've used before. 

Start by downloading the handout zip file for this problem. Then import into your blog database as follows:

mongoimport -d blog -c posts --drop posts.json
Now use the aggregation framework to calculate the author with the greatest number of comments.

To help you verify your work before submitting, the author with the fewest comments is Mariela Sherer and she commented 387 times. 

db.posts.aggregate([
 { "$project" : { "author" : "$comments.author"}}
,{ "$unwind"  : "$author"}
,{ "$group"   : { "_id"            : "$author",
   "numPosts" : { "$sum" : 1} }}
,{ "$sort" : { "numPosts" :-1}}
,{"$limit" : 1}
])

Please choose your answer below for the most prolific comment author:

Homework 5.2

Crunching the Zipcode dataset
Please calculate the average population of cities in California (abbreviation CA) and New York (NY) (taken together) with populations over 25,000. 

For this problem, assume that a city name that appears in more than one state represents two separate cities. 

Please round the answer to a whole number. 
Hint: The answer for CT and NJ (using this data set) is 38177. 

Please note:

  • Different states might have the same city name.
  • A city might have multiple zip codes.


For purposes of keeping the Hands On shell quick, we have used a subset of the data you previously used in zips.json, not the full set. This is why there are only 200 documents (and 200 zip codes), and all of them are in New York, Connecticut, New Jersey, and California. 

If you prefer, you may download the handout and perform your analysis on your machine with

> mongoimport -d test -c zips --drop small_zips.json


db.zips.aggregate([
 { "$match" : { "$or" : [ { "state" : "CA" },{ "state" :"NY" } ] }}

,{ "$group" : { "_id" : { "state" : "$state", "city" : "$city"}, 
                  "pop"  : { "$sum" : "$pop"}} }

,{ "$match" : { "pop" : { "$gt" : 25000 }}}

,{ "$group" : { "_id" : null,
"avg" : { "$avg" : "$pop"}} }
])

Once you've generated your aggregation query and found your answer, select it from the choices below. 

Homework 5.3

Who's the easiest grader on campus?
A set of grades are loaded into the grades collection. 

The documents look like this:

{
 "_id" : ObjectId("50b59cd75bed76f46522c392"),
 "student_id" : 10,
 "class_id" : 5,
 "scores" : [
  {
   "type" : "exam",
   "score" : 69.17634380939022
  },
  {
   "type" : "quiz",
   "score" : 61.20182926719762
  },
  {
   "type" : "homework",
   "score" : 73.3293624199466
  },
  {
   "type" : "homework",
   "score" : 15.206314042622903
  },
  {
   "type" : "homework",
   "score" : 36.75297723087603
  },
  {
   "type" : "homework",
   "score" : 64.42913107330241
  }
 ]
}
There are documents for each student (student_id) across a variety of classes (class_id). Note that not all students in the same class have the same exact number of assessments. Some students have three homework assignments, etc. 

Your task is to calculate the class with the best average student performance. This involves calculating an average for each student in each class of all non-quiz assessments and then averaging those numbers to get a class average. To be clear, each student's average includes only exams and homework grades. Don't include their quiz scores in the calculation. 

What is the class_id which has the highest average student perfomance? 

Hint/Strategy: You need to group twice to solve this problem. You must figure out the GPA that each student has achieved in a class and then average those numbers to get a class average. After that, you just need to sort. The class with the lowest average is the class with class_id=2. Those students achieved a class average of 37.6 

If you prefer, you may download the handout and perform your analysis on your machine with

> mongoimport -d test -c grades --drop grades.json

db.grades.aggregate([

 { "$unwind" : "$scores" }

,{ "$match"  : { "$or"    : [ { "scores.type" : "exam" },{ "scores.type" : "homework" }]} }

,{ "$group"  : { "_id"    : { "class" : "$class_id", "student" : "$student_id" },
       "stdAvg" : {"$avg" : "$scores.score" } } }

,{ "$group"  : { "_id"    : "$_id.class",
       "avg1"   : { "$avg" : "$stdAvg"}  } }

,{ "$project" : { "_id"   : false, "class": "$_id", "avg" : "$avg1" } }

,{ "$sort"   : { "avg"    : -1} }

,{ "$limit"  : 1}
]) 

Below, choose the class_id with the highest average student average.





Homework 5.4

Removing Rural Residents
In this problem you will calculate the number of people who live in a zip code in the US where the city starts with a digit. We will take that to mean they don't really live in a city. Once again, you will be using the zip code collection, which you will find in the 'handouts' link in this page. Import it into your mongod using the following command from the command line:
> mongoimport -d test -c zips --drop zips.json

If you imported it correctly, you can go to the test database in the mongo shell and conform that
> db.zips.count()

yields 29,467 documents. 

The project operator can extract the first digit from any field. For example, to extract the first digit from the city field, you could write this query:
db.zips.aggregate([
    {$project: 
     {
 first_char: {$substr : ["$city",0,1]},
     }  
   }
])
Using the aggregation framework, calculate the sum total of people who are living in a zip code where the city starts with a digit. Choose the answer below. 

Note that you will need to probably change your projection to send more info through than just that first character. Also, you will need a filtering step to get rid of all documents where the city does not start with a digital (0-9).

db.zips.aggregate([
 { "$project" : { "fc"   : { "$substr" : ["$city",0,1]},           
"pop"  : "$pop" }}
,{ "$match"   : { "fc" : /^[0123456789]/   }}
,{ "$group"   : { "_id" :  null,
   "pop1" : { "$sum" : "$pop"}}}
,{ "$project" : { "_id" : false, "pop" : "$pop1"}}
])


jueves, 4 de diciembre de 2014

MongoDB course for developers. unit 6/8. Application Engineering

Write Concern





It's how concern are your writes complete before you get the responses back ?
All of this is controlled by the drivers. The driver and mongo shell will execute for you the function getLastError after the write operation. In mongo the operations are not acknowledge and requires a second call to getLastError function.
Every single time yo do an insert o update (a single operation), the function getLastError is called in order to get a possible error. If you use a driver to access mongodb database you can decide if these drivers will call or not the function getLastError.
The function getLastError can have two parameters:
  • w: (write) if it's equal to 1, it determines weather or not you want to wait for the write operation to be acknowledge.
  • j: (journal = log to disk with the operations with the data) if it's equal to 1, getLastError waits until the journal commits to disk
Values of parameters:
  • w = 0, j = 0 : fire and forget 
  • w = 1, j = 0 : wait for a simple acknowlegement from mongo that receives the write. BY DEFAULT
  • w = 0 or 1 , j = 1 : wait for the write commits the journal. TO SURE THAT YOU ARRIVE TO DISK BEFORE YOU GO ON 
Provided you assume that the disk is persistent, what are the w and j settings required to guarantee that an insert or update has been written all the way to disk.


Network Errors

We can have network errors that provoque a write or insert cannot arrive to disk even though the options of write concern will set to true (w=1, j=1). We can never complete sure what exactly happened in the transaction. In order to solve this situations you can use a 'try exception' to get the error and offer a solution.

What are the reasons why an application may receive an error back even if the write was successful. Check all that apply.

The Pymongo Driver

The api web site is api.mongodb.org . This is the main directory of all the drivers to use with mongodb as a developer.

The current driver will always be at api.mongodb.org/python/current

The most important takeaways are that you should be using pymongo.MongoClient() to connect to a standalone server, or if you're connecting to a replica set, an even better option is pymongo.MongoReplicaSetClient().

Which of the following are valid, supported ways to connect to a server with pymongo?


Introduction to Replication

Replication offers fault tolerance in order to the system continues working when a node goes down or we have an accident like a fire.

The solution of mongo is building a replica set. A replica set is a group of nodes with mongod that works mirroring each other the data. The is one primary node and the others are secondaries dynamically.

The operation is that your application and its drivers stay connected to the primary node, and will write to the primary (you can only write to the primary). If the primary goes down, the reamining nodes will perfom an election to elect a new primary having a strict majority of the original nodes.

The minimun number of nodes to buid a replica set is three and you can have an arbiter node to decide which one will be primary in case of a tie.

What is the minimum original number of nodes needed to assure the election of a new Primary if a node goes down?


Replica Set Elections

The Replica set is totally transparent for applications that will continue working without any break.

Types of Replica set nodes:
  • Regular: has the data and it's the most normal type of node. It can be a primary or secondary
  • Arbiter: It can be a regular node. It's used for voting purposes. If you have even number of replica set nodes you need to make sure that there's an arbiter node in order to have a strict majority to elect a node as primary
  • Delayed/Regular: It offers the possibility to delay behind other nodes to recover data in a fast way. It cannot be a primary but can participate voting the election. Its priority is set to zero.
  • Hidden: it's often used for analytics.
Setting the priority to zero, the node will not elected as a primary. It cannot be a primary but can participate voting the election. Its priority is set to zero.
 
Wihich types of nodes can participate in elections of a new primary?

Write Consistency

The writes are sent to the primary node but the reads can send to any node (by default is set to the primary node to have a good consistency) but keep in mind tha the lag between any two nodes is no guaranteed becase the replication is asynchronous and data read cannot exist at the time of reading.
 
During the time when failover is occurring, can writes successfully complete?

Creating a Replica Set

In order to create a Replica set, we have to create and next initizalize their nodes (in this case the three nodes are in one server with differents ports. In other case it will be necessary specify th target host and port for each node):

1) Create the nodes of Replica set:
mongod --port 27017 --dbpath "/var/lib/mongodb/data/rs1" --replSet group1 --logpath "log_1.log" --oplogSize 200 --fork --smallfiles

mongod --port 27018 --dbpath "/var/lib/mongodb/data/rs2" --replSet group1 --logpath "log_2.log" --oplogSize 200 --fork --smallfiles

mongod --port 27019 --dbpath "/var/lib/mongodb/data/rs3" --replSet group1 --logpath "log_3.log" --oplogSize 200 --fork --smallfiles

2) Initialise the Replica set creating a variable within the configuration and loading it from the shell.

config = {
    "_id" : "group1",
    "members" : [
    //node 1
     { "_id" : 1 , "host" : "localhost:27017"}
    //node 2
    ,{ "_id" : 2 , "host" : "localhost:27018"}
    //node 3
    ,{ "_id" : 3 , "host" : "localhost:27019"}
    ]
};

rs.initiate(config);

To know the status of replica set:

rs.status();

To allow readings from an slave node it's necessary to execute in the secondary node:

rs.slaveOk()
 
To show which node is the master:

rs.isMaster()

To force the current replica set member to step down as primary and then attempt to avoid election as primary for the designated number of seconds (60 second by default). Produces an error if the current member is not the primary.

rs.stepDown()

To show help of replica set options:

rs.help()

Which command, when issued from the mongo shell, will allow you to read from a secondary?


Replica Set Internals

In the video how long did it take to elect a new primary?

 

Failover and Rollback

What happens if a node comes back up as a secondary after a period of being offline and the oplog has looped on the primary?

Connecting to a Replica Set from Pymongo

c = pymongo.MongoClient(host=
[
                 "mongodb://localhost:27017",
                 "mongodb://localhost:27018",
                 "mongodb://localhost:27019"
],
                 replicaSet="rs1",
                 w=1, j=True)
or
 
c = pymongo.MongoClient(host=
[
                 "mongodb://localhost:27017"
],
                 replicaSet="rs1",
                 w=1, j=True)
 
If you leave a replica set node out of the seedlist within the driver, what will happen?

What happens when the failover occurs

When the election is happening you can't complete writes or reads because yo don't have a primary to go to.

What will happen if the following statement is executed in Python during a primary election?
db.test.insert({'x':1})
 

Detecting Failover

If you catch exceptions during failover, are you guaranteed to have your writes succeed?
 

Proper Handling of Failover

Example in python of code to use in replica set to hand a failover:
 
def writesome():
    # let's do some inserting
    for i in range(0,1000000):
        for retries in range(0,3):
 
        # 'doc' is here because if a exception is produced the test.insert(doc) 
        # will insert a document with a different id and it will not provoke a  
        # duplicateKeyError

            doc = {'i':i}           
        try:
                test.insert(doc)
                print "Inserted " + str(i)
                break
            except pymongo.errors.DuplicateKeyError:
                print "Duplicate key error"
                break
            except:
                print sys.exc_info()[0]
                print "Retrying..."
                time.sleep(5)
        time.sleep(.5)

 
If this code guaranteed to get the write done if failover occurs:
doc = {'i':i}
        for retries in range(0,3):

            try:
                test.insert(doc)
                print "Inserted " + str(i)
                break
            except pymongo.errors.DuplicateKeyError:
                print "Duplicate key error"
                break
            except:
                print sys.exc_info()[0]
                print "Retrying..."
                time.sleep(5)


Write concern revisited

Options:
  • w = 1: it will wait until the primary node makes the write
  • w = 2: it will wait until two nodes make the write (if we have three nodes)
  • w = 3: it will wait until three nodes make the write (if we have three nodes)
  • j = 1: it will wait until the primary node 
  • wtimeout (seconds): how long you wiloing to wait for the writes to be acknowledged by the secondaries. (it can be set in the drivers)
 There are three differents places where can be set this options:
  1. on the connection
  2. on the connections inside the driver
  3. in the configuration itself of the Replica set you can set default values
 If you set w=1 and j=1, is it possible to wind up rolling back a committed write to the primary on failover?


Read preferences

You can specify read preference to read specific secondary nodes.
Options in Pymongo:
  1. always read on the primary
  2. always read on the secondary. If there isn't secondary the reads cannot do
  3. secondary preference and if there isn't it will read from a primary
  4. primary preference and if there isn't it will read from a secondary
  5. the nearest node
  6. by tagging. You can assing tags to nodes in order to name them
 You can configure your applications via the drivers to read from secondary nodes within a replica set. What are the reasons that you might not want to do that? Check all that apply.

Implications of replication

  1. Seed list to ensure that an election will done when the primary goes down
  2. Write concern: the idea of waiting for some number of nodes to acknowlege the writes through to w parameter, the j parameter which lets it wait or not for the primary node to commit that write to disk. An wtimeout parameter, which is how long you are going to wait to see that your write replicated to other members of the replica set. 
  3. Read preferences
  4. Errors can have: errors can always happen because of transient situations like failover occuring, or they can happen because there are network errors that occurs o errors in terms of violating the unique key constraints
To create a robust application it is necessary to check for exceptions of read and write operations to database in order to make sure that if anything comes up the application will know it. It is necessary to make sure we understand the application of what data has been committed and whata data is durable in the application.
  1. Seed list to ensure that an election will done when the primary goes down
  2. Write concern: the idea of waiting for some number of nodes to ack
If you set w=4 on a connection and there are only three nodes in the replica set, how long will you wait in PyMongo for a response from an insert if you don't set a timeout?

 

Introduction to Sharding

 This is an approach to horizontal scalability. Every shard node can have their replica set because of this we have a lot of hosts involved.

In order to distribute the data, mongo uses a router named 'mongos' that's going to take care of the distribution. It's going to keep some sort of connection pool or knowledge of all the different hosts, and it's going to route them properly.

shard_key: something that is going to determine, it's some part of the document itself  (the _id of the document)
 
Once you make the decission of what kind of shard key to use, mongo will then break the collection into chunks and decide what shard each of the chunks lives on a range-based way, and then any query that you make, which now has to be routed to a Mongo OS will then go to the appropiate shards to answer your query.
 
If the shard key is not include in a find operation and there are 3 shards, each one a replica set with 3 nodes, how many nodes will see the find operation?

Building a Sharded Environment

If you want to build a production system with two shards, each one a replica set with three nodes, how may mongod processes must you start?
2 shards has 6 nodes. 3 config nodes

Implications of Sharding

 Some things to remember in a shard environment:
  1. Every document needs to include the shard key
  2. The shard key is immutable: yo cannot change the shard key inside the document
  3. it needs an index that starts with the shard key but it cannot be a multiple index
  4. when do an update, it's necessary to specify the shard key or specify that multi is true
  5. no shard key means scatter gather operation, which could be expensive
  6. you can't have a unique key, no unique index, unless it's also part of the shard key
Suppose you wanted to shard the zip code collection after importing it. You want to shard on zip code. What index would be required to allow MongoDB to shard on zip code?

Sharding + Replication

Suppose you want to run multiple mongos routers for redundancy. What level of the stack will assure that you can failover to a different mongos from within your application?

Choosing a Shard Key

  1. Sufficient cardinality: in order to mongo can distribute the documents in shards
  2. avoid hotspots in writes: monotically increasing: the shard key will not provoque that all the writes will go to an specific shard
You are building a facebook competitor called footbook that will be a mobile social network of feet. You have decided that your primary data structure for posts to the wall will look like this:
{'username':'toeguy',
     'posttime':ISODate("2012-12-02T23:12:23Z"),
     "randomthought": "I am looking at my feet right now",
     'visible_to':['friends','family', 'walkers']}
Thinking about the tradeoffs of shard key selection, select the true statements below.