gunggum
내용
[subtitle-meta] lang=en auto=True
Everyone, AIAPI costs are
not just simple expenses, but a margin strategy. Let's start
with the first step, the prompt diet
. Simply
eliminating unnecessary words reduces token costs by up to 20%
.
Automatically remove low-importance words with a talk like LLML Lingu
. Please use the C-Max toe setting to
cut off the AI's rambling answers
. If you include too many examples, the
cost increases exponentially. The
correct answer is to reduce the example to one by specifying the Jason structure.
Rewrite costs with a second-stage architecture design.
Get a 50% discount on recurring upfront costs with prompt caching.
Always pin the registry document or system prompt to the top
. Sementi caching
completely blocks API calls for similar questions
.
Stores the question meaning of the vector DV and
returns the answer immediately. The key to the conversation history is to
summarize or send only the last 5
. Third step: You do not
need to use expensive models for every task
. Handle simple sorting
with a lightweight model that is twenty times cheaper.
We call the high-performance main model only for complex reasoning and coding.
Process nighttime data analysis using the Batch API to receive a 50% discount. For tasks that do not require real-time processing,
bundle them asynchronously and send them. In the long run, a
small model fine-tuned with our data
is the answer. Self-serving is much more advantageous in terms of security and cost
. You
can save 30% with immediate application and 70% with long-term optimization. You
can start using the prompt compression and batch APIs right away today. Cement
caching and fine tuning
shine brighter the larger the scale.
True efficiency comes from designing together with engineering and business
.
Please leave your service use cases in the comments. Please help us
make the next episode by subscribing and liking.
원문 자막 펼치기
WEBVTT Kind: captions Language: en 00:00:00.120 --> 00:00:02.030 align:start position:0% Everyone, <00:00:00.653><c>AIAPI </c><00:00:01.186><c>costs </c><00:00:01.719><c>are</c> 00:00:02.030 --> 00:00:02.040 align:start position:0% Everyone, AIAPI costs are 00:00:02.040 --> 00:00:04.550 align:start position:0% Everyone, AIAPI costs are not <00:00:02.297><c>just </c><00:00:02.554><c>simple </c><00:00:02.811><c>expenses, </c><00:00:03.068><c>but </c><00:00:03.325><c>a </c><00:00:03.582><c>margin </c><00:00:03.839><c>strategy. </c><00:00:04.096><c>Let's </c><00:00:04.353><c>start</c> 00:00:04.550 --> 00:00:04.560 align:start position:0% not just simple expenses, but a margin strategy. Let's start 00:00:04.560 --> 00:00:06.510 align:start position:0% not just simple expenses, but a margin strategy. Let's start with <00:00:04.773><c>the </c><00:00:04.986><c>first </c><00:00:05.199><c>step, </c><00:00:05.412><c>the </c><00:00:05.625><c>prompt </c><00:00:05.838><c>diet</c> 00:00:06.510 --> 00:00:06.520 align:start position:0% with the first step, the prompt diet 00:00:06.520 --> 00:00:09.030 align:start position:0% with the first step, the prompt diet . <00:00:08.679><c>Simply</c> 00:00:09.030 --> 00:00:09.040 align:start position:0% . Simply 00:00:09.040 --> 00:00:11.190 align:start position:0% . Simply eliminating <00:00:09.164><c>unnecessary </c><00:00:09.288><c>words </c><00:00:09.412><c>reduces </c><00:00:09.536><c>token </c><00:00:09.660><c>costs </c><00:00:09.784><c>by </c><00:00:09.908><c>up </c><00:00:10.032><c>to </c><00:00:10.156><c>20%</c> 00:00:11.190 --> 00:00:11.200 align:start position:0% eliminating unnecessary words reduces token costs by up to 20% 00:00:11.200 --> 00:00:13.509 align:start position:0% eliminating unnecessary words reduces token costs by up to 20% . 00:00:13.509 --> 00:00:13.519 align:start position:0% . 00:00:13.519 --> 00:00:15.110 align:start position:0% . Automatically <00:00:13.661><c>remove </c><00:00:13.803><c>low-importance </c><00:00:13.945><c>words </c><00:00:14.087><c>with </c><00:00:14.229><c>a </c><00:00:14.371><c>talk </c><00:00:14.513><c>like </c><00:00:14.655><c>LLML </c><00:00:14.797><c>Lingu</c> 00:00:15.110 --> 00:00:15.120 align:start position:0% Automatically remove low-importance words with a talk like LLML Lingu 00:00:15.120 --> 00:00:17.670 align:start position:0% Automatically remove low-importance words with a talk like LLML Lingu . <00:00:15.405><c>Please </c><00:00:15.690><c>use </c><00:00:15.975><c>the </c><00:00:16.260><c>C-Max </c><00:00:16.545><c>toe </c><00:00:16.830><c>setting </c><00:00:17.115><c>to</c> 00:00:17.670 --> 00:00:17.680 align:start position:0% . Please use the C-Max toe setting to 00:00:17.680 --> 00:00:19.550 align:start position:0% . Please use the C-Max toe setting to cut <00:00:18.008><c>off </c><00:00:18.336><c>the </c><00:00:18.664><c>AI's </c><00:00:18.992><c>rambling </c><00:00:19.320><c>answers</c> 00:00:19.550 --> 00:00:19.560 align:start position:0% cut off the AI's rambling answers 00:00:19.560 --> 00:00:21.990 align:start position:0% cut off the AI's rambling answers . <00:00:19.845><c>If </c><00:00:20.130><c>you </c><00:00:20.415><c>include </c><00:00:20.700><c>too </c><00:00:20.985><c>many </c><00:00:21.270><c>examples, </c><00:00:21.555><c>the</c> 00:00:21.990 --> 00:00:22.000 align:start position:0% . If you include too many examples, the 00:00:22.000 --> 00:00:24.390 align:start position:0% . If you include too many examples, the cost <00:00:22.413><c>increases </c><00:00:22.826><c>exponentially. </c><00:00:23.239><c>The</c> 00:00:24.390 --> 00:00:26.349 align:start position:0% cost increases exponentially. The 00:00:26.349 --> 00:00:26.359 align:start position:0% 00:00:26.359 --> 00:00:28.710 align:start position:0% correct <00:00:26.512><c>answer </c><00:00:26.665><c>is </c><00:00:26.818><c>to </c><00:00:26.971><c>reduce </c><00:00:27.124><c>the </c><00:00:27.277><c>example </c><00:00:27.430><c>to </c><00:00:27.583><c>one </c><00:00:27.736><c>by </c><00:00:27.889><c>specifying </c><00:00:28.042><c>the </c><00:00:28.195><c>Jason </c><00:00:28.348><c>structure.</c> 00:00:28.710 --> 00:00:28.720 align:start position:0% correct answer is to reduce the example to one by specifying the Jason structure. 00:00:28.720 --> 00:00:31.390 align:start position:0% correct answer is to reduce the example to one by specifying the Jason structure. Rewrite <00:00:29.000><c>costs </c><00:00:29.280><c>with </c><00:00:29.560><c>a </c><00:00:29.840><c>second-stage </c><00:00:30.120><c>architecture </c><00:00:30.400><c>design.</c> 00:00:31.390 --> 00:00:33.190 align:start position:0% Rewrite costs with a second-stage architecture design. 00:00:33.190 --> 00:00:33.200 align:start position:0% 00:00:33.200 --> 00:00:36.190 align:start position:0% Get <00:00:33.472><c>a </c><00:00:33.744><c>50% </c><00:00:34.016><c>discount </c><00:00:34.288><c>on </c><00:00:34.560><c>recurring </c><00:00:34.832><c>upfront </c><00:00:35.104><c>costs </c><00:00:35.376><c>with </c><00:00:35.648><c>prompt </c><00:00:35.920><c>caching.</c> 00:00:36.190 --> 00:00:36.200 align:start position:0% Get a 50% discount on recurring upfront costs with prompt caching. 00:00:36.200 --> 00:00:38.549 align:start position:0% Get a 50% discount on recurring upfront costs with prompt caching. Always <00:00:36.396><c>pin </c><00:00:36.592><c>the </c><00:00:36.788><c>registry </c><00:00:36.984><c>document </c><00:00:37.180><c>or </c><00:00:37.376><c>system </c><00:00:37.572><c>prompt </c><00:00:37.768><c>to </c><00:00:37.964><c>the </c><00:00:38.160><c>top</c> 00:00:38.549 --> 00:00:38.559 align:start position:0% Always pin the registry document or system prompt to the top 00:00:38.559 --> 00:00:41.150 align:start position:0% Always pin the registry document or system prompt to the top . <00:00:39.679><c>Sementi </c><00:00:40.799><c>caching</c> 00:00:41.150 --> 00:00:41.160 align:start position:0% . Sementi caching 00:00:41.160 --> 00:00:42.630 align:start position:0% . Sementi caching completely <00:00:41.360><c>blocks </c><00:00:41.560><c>API </c><00:00:41.760><c>calls </c><00:00:41.960><c>for </c><00:00:42.160><c>similar </c><00:00:42.360><c>questions</c> 00:00:42.630 --> 00:00:42.640 align:start position:0% completely blocks API calls for similar questions 00:00:42.640 --> 00:00:44.270 align:start position:0% completely blocks API calls for similar questions . 00:00:44.270 --> 00:00:44.280 align:start position:0% . 00:00:44.280 --> 00:00:46.430 align:start position:0% . Stores <00:00:44.515><c>the </c><00:00:44.750><c>question </c><00:00:44.985><c>meaning </c><00:00:45.220><c>of </c><00:00:45.455><c>the </c><00:00:45.690><c>vector </c><00:00:45.925><c>DV </c><00:00:46.160><c>and</c> 00:00:46.430 --> 00:00:46.440 align:start position:0% Stores the question meaning of the vector DV and 00:00:46.440 --> 00:00:48.910 align:start position:0% Stores the question meaning of the vector DV and returns <00:00:46.629><c>the </c><00:00:46.818><c>answer </c><00:00:47.007><c>immediately. </c><00:00:47.196><c>The </c><00:00:47.385><c>key </c><00:00:47.574><c>to </c><00:00:47.763><c>the </c><00:00:47.952><c>conversation </c><00:00:48.141><c>history </c><00:00:48.330><c>is </c><00:00:48.519><c>to</c> 00:00:48.910 --> 00:00:48.920 align:start position:0% returns the answer immediately. The key to the conversation history is to 00:00:48.920 --> 00:00:50.910 align:start position:0% returns the answer immediately. The key to the conversation history is to summarize <00:00:49.153><c>or </c><00:00:49.386><c>send </c><00:00:49.619><c>only </c><00:00:49.852><c>the </c><00:00:50.085><c>last </c><00:00:50.318><c>5</c> 00:00:50.910 --> 00:00:50.920 align:start position:0% summarize or send only the last 5 00:00:50.920 --> 00:00:53.510 align:start position:0% summarize or send only the last 5 . <00:00:51.383><c>Third </c><00:00:51.846><c>step: </c><00:00:52.309><c>You </c><00:00:52.772><c>do </c><00:00:53.235><c>not</c> 00:00:53.510 --> 00:00:53.520 align:start position:0% . Third step: You do not 00:00:53.520 --> 00:00:55.150 align:start position:0% . Third step: You do not need <00:00:53.708><c>to </c><00:00:53.896><c>use </c><00:00:54.084><c>expensive </c><00:00:54.272><c>models </c><00:00:54.460><c>for </c><00:00:54.648><c>every </c><00:00:54.836><c>task</c> 00:00:55.150 --> 00:00:55.160 align:start position:0% need to use expensive models for every task 00:00:55.160 --> 00:00:57.709 align:start position:0% need to use expensive models for every task . <00:00:55.866><c>Handle </c><00:00:56.572><c>simple </c><00:00:57.278><c>sorting</c> 00:00:57.709 --> 00:00:57.719 align:start position:0% . Handle simple sorting 00:00:57.719 --> 00:01:00.270 align:start position:0% . Handle simple sorting with <00:00:57.874><c>a </c><00:00:58.029><c>lightweight </c><00:00:58.184><c>model </c><00:00:58.339><c>that </c><00:00:58.494><c>is </c><00:00:58.649><c>twenty </c><00:00:58.804><c>times </c><00:00:58.959><c>cheaper.</c> 00:01:00.270 --> 00:01:02.430 align:start position:0% with a lightweight model that is twenty times cheaper. 00:01:02.430 --> 00:01:02.440 align:start position:0% 00:01:02.440 --> 00:01:04.710 align:start position:0% We <00:01:02.472><c>call </c><00:01:02.504><c>the </c><00:01:02.536><c>high-performance </c><00:01:02.568><c>main </c><00:01:02.600><c>model </c><00:01:02.632><c>only </c><00:01:02.664><c>for </c><00:01:02.696><c>complex </c><00:01:02.728><c>reasoning </c><00:01:02.760><c>and </c><00:01:02.792><c>coding.</c> 00:01:04.710 --> 00:01:06.590 align:start position:0% We call the high-performance main model only for complex reasoning and coding. 00:01:06.590 --> 00:01:06.600 align:start position:0% 00:01:06.600 --> 00:01:09.429 align:start position:0% Process <00:01:06.718><c>nighttime </c><00:01:06.836><c>data </c><00:01:06.954><c>analysis </c><00:01:07.072><c>using </c><00:01:07.190><c>the </c><00:01:07.308><c>Batch </c><00:01:07.426><c>API </c><00:01:07.544><c>to </c><00:01:07.662><c>receive </c><00:01:07.780><c>a </c><00:01:07.898><c>50% </c><00:01:08.016><c>discount. </c><00:01:08.134><c>For </c><00:01:08.252><c>tasks </c><00:01:08.370><c>that </c><00:01:08.488><c>do </c><00:01:08.606><c>not </c><00:01:08.724><c>require </c><00:01:08.842><c>real-time </c><00:01:08.960><c>processing,</c> 00:01:09.429 --> 00:01:11.270 align:start position:0% Process nighttime data analysis using the Batch API to receive a 50% discount. For tasks that do not require real-time processing, 00:01:11.270 --> 00:01:11.280 align:start position:0% 00:01:11.280 --> 00:01:13.510 align:start position:0% bundle <00:01:11.484><c>them </c><00:01:11.688><c>asynchronously </c><00:01:11.892><c>and </c><00:01:12.096><c>send </c><00:01:12.300><c>them. </c><00:01:12.504><c>In </c><00:01:12.708><c>the </c><00:01:12.912><c>long </c><00:01:13.116><c>run, </c><00:01:13.320><c>a</c> 00:01:13.510 --> 00:01:13.520 align:start position:0% bundle them asynchronously and send them. In the long run, a 00:01:13.520 --> 00:01:15.310 align:start position:0% bundle them asynchronously and send them. In the long run, a small <00:01:13.808><c>model </c><00:01:14.096><c>fine-tuned </c><00:01:14.384><c>with </c><00:01:14.672><c>our </c><00:01:14.960><c>data</c> 00:01:15.310 --> 00:01:15.320 align:start position:0% small model fine-tuned with our data 00:01:15.320 --> 00:01:18.270 align:start position:0% small model fine-tuned with our data is <00:01:15.526><c>the </c><00:01:15.732><c>answer. </c><00:01:15.938><c>Self-serving </c><00:01:16.144><c>is </c><00:01:16.350><c>much </c><00:01:16.556><c>more </c><00:01:16.762><c>advantageous </c><00:01:16.968><c>in </c><00:01:17.174><c>terms </c><00:01:17.380><c>of </c><00:01:17.586><c>security </c><00:01:17.792><c>and </c><00:01:17.998><c>cost</c> 00:01:18.270 --> 00:01:18.280 align:start position:0% is the answer. Self-serving is much more advantageous in terms of security and cost 00:01:18.280 --> 00:01:20.830 align:start position:0% is the answer. Self-serving is much more advantageous in terms of security and cost . <00:01:18.920><c>You</c> 00:01:20.830 --> 00:01:22.710 align:start position:0% . You 00:01:22.710 --> 00:01:22.720 align:start position:0% 00:01:22.720 --> 00:01:25.670 align:start position:0% can <00:01:22.872><c>save </c><00:01:23.024><c>30% </c><00:01:23.176><c>with </c><00:01:23.328><c>immediate </c><00:01:23.480><c>application </c><00:01:23.632><c>and </c><00:01:23.784><c>70% </c><00:01:23.936><c>with </c><00:01:24.088><c>long-term </c><00:01:24.240><c>optimization. </c><00:01:24.392><c>You</c> 00:01:25.670 --> 00:01:27.870 align:start position:0% can save 30% with immediate application and 70% with long-term optimization. You 00:01:27.870 --> 00:01:27.880 align:start position:0% 00:01:27.880 --> 00:01:30.149 align:start position:0% can <00:01:28.033><c>start </c><00:01:28.186><c>using </c><00:01:28.339><c>the </c><00:01:28.492><c>prompt </c><00:01:28.645><c>compression </c><00:01:28.798><c>and </c><00:01:28.951><c>batch </c><00:01:29.104><c>APIs </c><00:01:29.257><c>right </c><00:01:29.410><c>away </c><00:01:29.563><c>today. </c><00:01:29.716><c>Cement</c> 00:01:30.149 --> 00:01:30.159 align:start position:0% can start using the prompt compression and batch APIs right away today. Cement 00:01:30.159 --> 00:01:32.190 align:start position:0% can start using the prompt compression and batch APIs right away today. Cement caching <00:01:30.666><c>and </c><00:01:31.173><c>fine </c><00:01:31.680><c>tuning</c> 00:01:32.190 --> 00:01:32.200 align:start position:0% caching and fine tuning 00:01:32.200 --> 00:01:34.510 align:start position:0% caching and fine tuning shine <00:01:32.504><c>brighter </c><00:01:32.808><c>the </c><00:01:33.112><c>larger </c><00:01:33.416><c>the </c><00:01:33.720><c>scale.</c> 00:01:34.510 --> 00:01:34.520 align:start position:0% shine brighter the larger the scale. 00:01:34.520 --> 00:01:36.990 align:start position:0% shine brighter the larger the scale. True <00:01:34.760><c>efficiency </c><00:01:35.000><c>comes </c><00:01:35.240><c>from </c><00:01:35.480><c>designing </c><00:01:35.720><c>together </c><00:01:35.960><c>with </c><00:01:36.200><c>engineering </c><00:01:36.440><c>and </c><00:01:36.680><c>business</c> 00:01:36.990 --> 00:01:37.000 align:start position:0% True efficiency comes from designing together with engineering and business 00:01:37.000 --> 00:01:38.950 align:start position:0% True efficiency comes from designing together with engineering and business . 00:01:38.950 --> 00:01:38.960 align:start position:0% . 00:01:38.960 --> 00:01:41.670 align:start position:0% . Please <00:01:39.080><c>leave </c><00:01:39.200><c>your </c><00:01:39.320><c>service </c><00:01:39.440><c>use </c><00:01:39.560><c>cases </c><00:01:39.680><c>in </c><00:01:39.800><c>the </c><00:01:39.920><c>comments. </c><00:01:40.040><c>Please </c><00:01:40.160><c>help </c><00:01:40.280><c>us</c> 00:01:41.670 --> 00:01:43.670 align:start position:0% Please leave your service use cases in the comments. Please help us 00:01:43.670 --> 00:01:43.680 align:start position:0% 00:01:43.680 --> 00:01:46.159 align:start position:0% make <00:01:43.725><c>the </c><00:01:43.770><c>next </c><00:01:43.815><c>episode </c><00:01:43.860><c>by </c><00:01:43.905><c>subscribing </c><00:01:43.950><c>and </c><00:01:43.995><c>liking.</c>