Buscar
Mostrando entradas con la etiqueta mongoDB_4_dev. Mostrar todas las entradas
Mostrando entradas con la etiqueta mongoDB_4_dev. Mostrar todas las entradas
miércoles, 18 de febrero de 2015
martes, 13 de enero de 2015
MongoDB course for developers. unit 5/8. Aggregation framework. Homeworks
Homework 5.1
Finding the most frequent author of comments on your blog
In this assignment you will use the aggregation framework to find the most frequent author of comments on your blog. We will be using a data set similar to ones we've used before.
Start by downloading the handout zip file for this problem. Then import into your blog database as follows:
mongoimport -d blog -c posts --drop posts.json
Now use the aggregation framework to calculate the author with the greatest number of comments.
To help you verify your work before submitting, the author with the fewest comments is Mariela Sherer and she commented 387 times.
db.posts.aggregate([
{ "$project" : { "author" : "$comments.author"}}
,{ "$unwind" : "$author"}
,{ "$group" : { "_id" : "$author",
"numPosts" : { "$sum" : 1} }}
,{ "$sort" : { "numPosts" :-1}}
,{"$limit" : 1}
])
Please choose your answer below for the most prolific comment author:
Finding the most frequent author of comments on your blog
In this assignment you will use the aggregation framework to find the most frequent author of comments on your blog. We will be using a data set similar to ones we've used before.
Start by downloading the handout zip file for this problem. Then import into your blog database as follows:
mongoimport -d blog -c posts --drop posts.jsonNow use the aggregation framework to calculate the author with the greatest number of comments.
To help you verify your work before submitting, the author with the fewest comments is Mariela Sherer and she commented 387 times.
db.posts.aggregate([
{ "$project" : { "author" : "$comments.author"}}
,{ "$unwind" : "$author"}
,{ "$group" : { "_id" : "$author",
"numPosts" : { "$sum" : 1} }}
,{ "$sort" : { "numPosts" :-1}}
,{"$limit" : 1}
])
Please choose your answer below for the most prolific comment author:
Homework 5.2
Crunching the Zipcode dataset
Please calculate the average population of cities in California (abbreviation CA) and New York (NY) (taken together) with populations over 25,000.
For this problem, assume that a city name that appears in more than one state represents two separate cities.
Please round the answer to a whole number.
Hint: The answer for CT and NJ (using this data set) is 38177.
Please note:
- Different states might have the same city name.
- A city might have multiple zip codes.
For purposes of keeping the Hands On shell quick, we have used a subset of the data you previously used in zips.json, not the full set. This is why there are only 200 documents (and 200 zip codes), and all of them are in New York, Connecticut, New Jersey, and California.
If you prefer, you may download the handout and perform your analysis on your machine with
> mongoimport -d test -c zips --drop small_zips.json
db.zips.aggregate([
{ "$match" : { "$or" : [ { "state" : "CA" },{ "state" :"NY" } ] }}
,{ "$group" : { "_id" : { "state" : "$state", "city" : "$city"},
"pop" : { "$sum" : "$pop"}} }
,{ "$match" : { "pop" : { "$gt" : 25000 }}}
,{ "$group" : { "_id" : null,
"avg" : { "$avg" : "$pop"}} }
])
Once you've generated your aggregation query and found your answer, select it from the choices below.
Crunching the Zipcode dataset
Please calculate the average population of cities in California (abbreviation CA) and New York (NY) (taken together) with populations over 25,000.
For this problem, assume that a city name that appears in more than one state represents two separate cities.
Please round the answer to a whole number.
Hint: The answer for CT and NJ (using this data set) is 38177.
Please note:
For purposes of keeping the Hands On shell quick, we have used a subset of the data you previously used in zips.json, not the full set. This is why there are only 200 documents (and 200 zip codes), and all of them are in New York, Connecticut, New Jersey, and California.
If you prefer, you may download the handout and perform your analysis on your machine with
,{ "$group" : { "_id" : { "state" : "$state", "city" : "$city"},
,{ "$match" : { "pop" : { "$gt" : 25000 }}}
,{ "$group" : { "_id" : null,
Once you've generated your aggregation query and found your answer, select it from the choices below.
Please calculate the average population of cities in California (abbreviation CA) and New York (NY) (taken together) with populations over 25,000.
For this problem, assume that a city name that appears in more than one state represents two separate cities.
Please round the answer to a whole number.
Hint: The answer for CT and NJ (using this data set) is 38177.
Please note:
- Different states might have the same city name.
- A city might have multiple zip codes.
For purposes of keeping the Hands On shell quick, we have used a subset of the data you previously used in zips.json, not the full set. This is why there are only 200 documents (and 200 zip codes), and all of them are in New York, Connecticut, New Jersey, and California.
If you prefer, you may download the handout and perform your analysis on your machine with
> mongoimport -d test -c zips --drop small_zips.json
db.zips.aggregate([
{ "$match" : { "$or" : [ { "state" : "CA" },{ "state" :"NY" } ] }}
,{ "$group" : { "_id" : { "state" : "$state", "city" : "$city"},
"pop" : { "$sum" : "$pop"}} }
,{ "$match" : { "pop" : { "$gt" : 25000 }}}
,{ "$group" : { "_id" : null,
"avg" : { "$avg" : "$pop"}} }
])
Homework 5.3
Who's the easiest grader on campus?
A set of grades are loaded into the grades collection.
The documents look like this:
{
"_id" : ObjectId("50b59cd75bed76f46522c392"),
"student_id" : 10,
"class_id" : 5,
"scores" : [
{
"type" : "exam",
"score" : 69.17634380939022
},
{
"type" : "quiz",
"score" : 61.20182926719762
},
{
"type" : "homework",
"score" : 73.3293624199466
},
{
"type" : "homework",
"score" : 15.206314042622903
},
{
"type" : "homework",
"score" : 36.75297723087603
},
{
"type" : "homework",
"score" : 64.42913107330241
}
]
}
There are documents for each student (student_id) across a variety of classes (class_id). Note that not all students in the same class have the same exact number of assessments. Some students have three homework assignments, etc.
Your task is to calculate the class with the best average student performance. This involves calculating an average for each student in each class of all non-quiz assessments and then averaging those numbers to get a class average. To be clear, each student's average includes only exams and homework grades. Don't include their quiz scores in the calculation.
What is the class_id which has the highest average student perfomance?
Hint/Strategy: You need to group twice to solve this problem. You must figure out the GPA that each student has achieved in a class and then average those numbers to get a class average. After that, you just need to sort. The class with the lowest average is the class with class_id=2. Those students achieved a class average of 37.6
If you prefer, you may download the handout and perform your analysis on your machine with
> mongoimport -d test -c grades --drop grades.json
db.grades.aggregate([
Who's the easiest grader on campus?
A set of grades are loaded into the grades collection.
The documents look like this:
Your task is to calculate the class with the best average student performance. This involves calculating an average for each student in each class of all non-quiz assessments and then averaging those numbers to get a class average. To be clear, each student's average includes only exams and homework grades. Don't include their quiz scores in the calculation.
What is the class_id which has the highest average student perfomance?
Hint/Strategy: You need to group twice to solve this problem. You must figure out the GPA that each student has achieved in a class and then average those numbers to get a class average. After that, you just need to sort. The class with the lowest average is the class with class_id=2. Those students achieved a class average of 37.6
If you prefer, you may download the handout and perform your analysis on your machine with
db.grades.aggregate([
A set of grades are loaded into the grades collection.
The documents look like this:
{
"_id" : ObjectId("50b59cd75bed76f46522c392"),
"student_id" : 10,
"class_id" : 5,
"scores" : [
{
"type" : "exam",
"score" : 69.17634380939022
},
{
"type" : "quiz",
"score" : 61.20182926719762
},
{
"type" : "homework",
"score" : 73.3293624199466
},
{
"type" : "homework",
"score" : 15.206314042622903
},
{
"type" : "homework",
"score" : 36.75297723087603
},
{
"type" : "homework",
"score" : 64.42913107330241
}
]
}
There are documents for each student (student_id) across a variety of classes (class_id). Note that not all students in the same class have the same exact number of assessments. Some students have three homework assignments, etc. Your task is to calculate the class with the best average student performance. This involves calculating an average for each student in each class of all non-quiz assessments and then averaging those numbers to get a class average. To be clear, each student's average includes only exams and homework grades. Don't include their quiz scores in the calculation.
What is the class_id which has the highest average student perfomance?
Hint/Strategy: You need to group twice to solve this problem. You must figure out the GPA that each student has achieved in a class and then average those numbers to get a class average. After that, you just need to sort. The class with the lowest average is the class with class_id=2. Those students achieved a class average of 37.6
If you prefer, you may download the handout and perform your analysis on your machine with
> mongoimport -d test -c grades --drop grades.json
db.grades.aggregate([
{ "$unwind" : "$scores" }
,{ "$match" : { "$or" : [ { "scores.type" : "exam" },{ "scores.type" : "homework" }]} }
,{ "$group" : { "_id" : { "class" : "$class_id", "student" : "$student_id" },
"stdAvg" : {"$avg" : "$scores.score" } } }
,{ "$group" : { "_id" : "$_id.class",
"avg1" : { "$avg" : "$stdAvg"} } }
,{ "$project" : { "_id" : false, "class": "$_id", "avg" : "$avg1" } }
,{ "$sort" : { "avg" : -1} }
,{ "$limit" : 1}
])
Below, choose the class_id with the highest average student average.
Homework 5.4
Removing Rural Residents
In this problem you will calculate the number of people who live in a zip code in the US where the city starts with a digit. We will take that to mean they don't really live in a city. Once again, you will be using the zip code collection, which you will find in the 'handouts' link in this page. Import it into your mongod using the following command from the command line:
> mongoimport -d test -c zips --drop zips.json
If you imported it correctly, you can go to the test database in the mongo shell and conform that
> db.zips.count()
yields 29,467 documents.
The project operator can extract the first digit from any field. For example, to extract the first digit from the city field, you could write this query:
db.zips.aggregate([
{$project:
{
first_char: {$substr : ["$city",0,1]},
}
}
])
Using the aggregation framework, calculate the sum total of people who are living in a zip code where the city starts with a digit. Choose the answer below.
Note that you will need to probably change your projection to send more info through than just that first character. Also, you will need a filtering step to get rid of all documents where the city does not start with a digital (0-9).
Removing Rural Residents
In this problem you will calculate the number of people who live in a zip code in the US where the city starts with a digit. We will take that to mean they don't really live in a city. Once again, you will be using the zip code collection, which you will find in the 'handouts' link in this page. Import it into your mongod using the following command from the command line:
> mongoimport -d test -c zips --drop zips.json
If you imported it correctly, you can go to the test database in the mongo shell and conform that
> db.zips.count()
yields 29,467 documents.
The project operator can extract the first digit from any field. For example, to extract the first digit from the city field, you could write this query:
db.zips.aggregate([
{$project:
{
first_char: {$substr : ["$city",0,1]},
}
}
])
Using the aggregation framework, calculate the sum total of people who are living in a zip code where the city starts with a digit. Choose the answer below.
Note that you will need to probably change your projection to send more info through than just that first character. Also, you will need a filtering step to get rid of all documents where the city does not start with a digital (0-9).
jueves, 4 de diciembre de 2014
MongoDB course for developers. unit 6/8. Application Engineering
Write Concern
- w: (write) if it's equal to 1, it determines weather or not you want to wait for the write operation to be acknowledge.
- j: (journal = log to disk with the operations with the data) if it's equal to 1, getLastError waits until the journal commits to disk
- w = 0, j = 0 : fire and forget
- w = 1, j = 0 : wait for a simple acknowlegement from mongo that receives the write. BY DEFAULT
- w = 0 or 1 , j = 1 : wait for the write commits the journal. TO SURE THAT YOU ARRIVE TO DISK BEFORE YOU GO ON
Network Errors
The Pymongo Driver
The current driver will always be at api.mongodb.org/python/current
The most important takeaways are that you should be using pymongo.MongoClient() to connect to a standalone server, or if you're connecting to a replica set, an even better option is pymongo.MongoReplicaSetClient().
Which of the following are valid, supported ways to connect to a server with pymongo?
Introduction to Replication
The solution of mongo is building a replica set. A replica set is a group of nodes with mongod that works mirroring each other the data. The is one primary node and the others are secondaries dynamically.
The operation is that your application and its drivers stay connected to the primary node, and will write to the primary (you can only write to the primary). If the primary goes down, the reamining nodes will perfom an election to elect a new primary having a strict majority of the original nodes.
The minimun number of nodes to buid a replica set is three and you can have an arbiter node to decide which one will be primary in case of a tie.
Replica Set Elections
Types of Replica set nodes:
- Regular: has the data and it's the most normal type of node. It can be a primary or secondary
- Arbiter: It can be a regular node. It's used for voting purposes. If you have even number of replica set nodes you need to make sure that there's an arbiter node in order to have a strict majority to elect a node as primary
- Delayed/Regular: It offers the possibility to delay behind other nodes to recover data in a fast way. It cannot be a primary but can participate voting the election. Its priority is set to zero.
- Hidden: it's often used for analytics.
Write Consistency
Creating a Replica Set
1) Create the nodes of Replica set:
mongod --port 27017 --dbpath "/var/lib/mongodb/data/rs1" --replSet group1 --logpath "log_1.log" --oplogSize 200 --fork --smallfiles
mongod --port 27018 --dbpath "/var/lib/mongodb/data/rs2" --replSet group1 --logpath "log_2.log" --oplogSize 200 --fork --smallfiles
mongod --port 27019 --dbpath "/var/lib/mongodb/data/rs3" --replSet group1 --logpath "log_3.log" --oplogSize 200 --fork --smallfiles
2) Initialise the Replica set creating a variable within the configuration and loading it from the shell.
config = {
"_id" : "group1",
"members" : [
//node 1
{ "_id" : 1 , "host" : "localhost:27017"}
//node 2
,{ "_id" : 2 , "host" : "localhost:27018"}
//node 3
,{ "_id" : 3 , "host" : "localhost:27019"}
]
};
rs.initiate(config);
To know the status of replica set:
rs.status();
To allow readings from an slave node it's necessary to execute in the secondary node:
rs.slaveOk()
rs.isMaster()
rs.stepDown()
rs.help()
Replica Set Internals
Failover and Rollback
What happens if a node comes back up as a secondary after a period of being offline and the oplog has looped on the primary?
Connecting to a Replica Set from Pymongo
"mongodb://localhost:27018",
"mongodb://localhost:27019"
replicaSet="rs1",
w=1, j=True)
replicaSet="rs1",
w=1, j=True)
What happens when the failover occurs
db.test.insert({'x':1})
Detecting Failover
Proper Handling of Failover
# let's do some inserting
for i in range(0,1000000):
for retries in range(0,3):
doc = {'i':i}
try:
test.insert(doc)
print "Inserted " + str(i)
break
except pymongo.errors.DuplicateKeyError:
print "Duplicate key error"
break
except:
print sys.exc_info()[0]
print "Retrying..."
time.sleep(5)
time.sleep(.5)
doc = {'i':i}
for retries in range(0,3):
try:
test.insert(doc)
print "Inserted " + str(i)
break
except pymongo.errors.DuplicateKeyError:
print "Duplicate key error"
break
except:
print sys.exc_info()[0]
print "Retrying..."
time.sleep(5)
Write concern revisited
- w = 1: it will wait until the primary node makes the write
- w = 2: it will wait until two nodes make the write (if we have three nodes)
- w = 3: it will wait until three nodes make the write (if we have three nodes)
- j = 1: it will wait until the primary node
- wtimeout (seconds): how long you wiloing to wait for the writes to be acknowledged by the secondaries. (it can be set in the drivers)
- on the connection
- on the connections inside the driver
- in the configuration itself of the Replica set you can set default values
Read preferences
- always read on the primary
- always read on the secondary. If there isn't secondary the reads cannot do
- secondary preference and if there isn't it will read from a primary
- primary preference and if there isn't it will read from a secondary
- the nearest node
- by tagging. You can assing tags to nodes in order to name them
Implications of replication
- Seed list to ensure that an election will done when the primary goes down
- Write concern: the idea of waiting for some number of nodes to acknowlege the writes through to w parameter, the j parameter which lets it wait or not for the primary node to commit that write to disk. An wtimeout parameter, which is how long you are going to wait to see that your write replicated to other members of the replica set.
- Read preferences
- Errors can have: errors can always happen because of transient situations like failover occuring, or they can happen because there are network errors that occurs o errors in terms of violating the unique key constraints
- Seed list to ensure that an election will done when the primary goes down
- Write concern: the idea of waiting for some number of nodes to ack
Introduction to Sharding
Building a Sharded Environment
Implications of Sharding
- Every document needs to include the shard key
- The shard key is immutable: yo cannot change the shard key inside the document
- it needs an index that starts with the shard key but it cannot be a multiple index
- when do an update, it's necessary to specify the shard key or specify that multi is true
- no shard key means scatter gather operation, which could be expensive
- you can't have a unique key, no unique index, unless it's also part of the shard key
Sharding + Replication
Choosing a Shard Key
- Sufficient cardinality: in order to mongo can distribute the documents in shards
- avoid hotspots in writes: monotically increasing: the shard key will not provoque that all the writes will go to an specific shard
{'username':'toeguy',
'posttime':ISODate("2012-12-02T23:12:23Z"),
"randomthought": "I am looking at my feet right now",
'visible_to':['friends','family', 'walkers']}
Thinking about the tradeoffs of shard key selection, select the true statements below.
Suscribirse a:
Entradas (Atom)