Topic 1 Question 35
A Data Science team is designing a dataset repository where it will store a large amount of training data commonly used in its machine learning models. As Data Scientists may create an arbitrary number of new datasets every day, the solution has to scale automatically and be cost-effective. Also, it must be possible to explore the data using SQL. Which storage scheme is MOST adapted to this scenario?
Store datasets as files in Amazon S3.
Store datasets as files in an Amazon EBS volume attached to an Amazon EC2 instance.
Store datasets as tables in a multi-node Amazon Redshift cluster.
Store datasets as global tables in Amazon DynamoDB.
ユーザの投票
コメント(8)
Ans: A (S3) is most cost effective
👍 15rsimham2021/09/22A : S3 cost effective + athena ( not c redshift dont support unstructured data)
👍 7sonalev4192021/11/01I would say C https://docs.aws.amazon.com/redshift/latest/mgmt/working-with-clusters.html "For workloads that require ever-growing storage, managed storage lets you automatically scale your data warehouse storage capacity without adding and paying for additional nodes."
👍 3syu31svc2021/10/05
シャッフルモード